All work

Rail ETA

Predicts how late a train will reach every station still ahead.

Team project, Smart India Hackathon

Code

Our team's answer to Smart India Hackathon problem SIH26028, Dynamic Forecast of ETA for Coaching Trains. Indian Railways copies the current delay forward. Our model, Aagaman, forecasts the delay at each station ahead.

The data

  • A collector on a home server kept sweeping live running status from NTES and storing every observation.
  • The final test used 12,509 complete journeys of 2,180 trains over 8 days, the full window NTES keeps.
  • That came to 664,461 forecasts scored against what actually happened.

The model

  • Gradient boosting with scikit-learn, on 21 features such as current delay, distance to the target station, weather forecast and traffic ahead.
  • Distance replaced stop count, because one stop is about 250 km on a Rajdhani and about 30 km on a Shatabdi.

The result

  • Mean error of 20.83 minutes, against 21.68 for the rule in use today: 3.9% better.
  • On 435 trains it never saw in training, 2.1% better.

The hard part

Predicting the arrival delay directly broke on very late trains: a tree model cannot go past the values it was trained on.

It predicts the delay added between here and the target station instead, and adds it to the delay the train already has.

Mean error of 20.83 minutes, against 21.68 for the rule in use today.

Built with

  • Python
  • scikit-learn
  • SQLite
  • systemd on a home server
Chart for the 12951 Mumbai Rajdhani, 22 minutes late at Borivali in fog season. Today's method stays flat at 22 minutes; Aagaman rises to 51 minutes by New Delhi.
One real forecast from our slides: 29 minutes apart by New Delhi.
Results table: 2,180 trains, 12,509 journeys, 664,461 forecasts, 3.9% better than the Indian Railways rule.
Measured on real trains, from our slides.

That is Rail ETA.

Read to the end, and its shape is yours.