CS 506 Final Project
Video: https://youtu.be/BJXz8D5m6j4
Our project aims to predict when the Green Line will have reliability issues and identify crowding hotspots, so students can plan their commutes better.
The Problem: Students at specific stops and times significantly contribute to MBTA train delays. We wanted to predict surges in passenger volume on the MBTA Green Line and create a tool that students can use while commuting.
The Goal: Identify transit crowd hotspots with high accuracy and visualize them so students can avoid delays and plan when to take the T.
-
Install dependencies:
make install # Or manually: pip install -r requirements.txt -
Run the full pipeline:
make all
This will: install dependencies → download/process data → run tests → generate visualizations
506MBTAProject/
├── src/ # Source code
│ ├── mbta/ # MBTA data processing
│ │ ├── class_schedules.py # BU class schedule patterns
│ │ ├── historical.py # Reliability data processing
│ │ ├── weather.py # Weather data processing
│ │ ├── merging.py # Data merging pipeline
│ │ ├── mbta.py # MBTA API utilities
│ │ └── stop_names.py # Stop ID to name mapping
│ ├── models/ # Modeling code
│ │ ├── feature_engineering.py # Feature creation (119 features)
│ │ └── model.py # Final model (Random Forest)
│ └── integration/ # Data integration
│ ├── lamp_data_integration.py # LAMP performance data
│ └── lamp_alerts_integration.py # LAMP alerts data
├── scripts/ # Utility scripts
│ └── download_more_lamp_data.py
├── notebooks/ # Analysis notebooks
│ └── analysis.ipynb
├── visualization/ # Interactive visualizations
│ └── mbta_map_viz.py
├── tests/ # Test files
├── output/ # Generated visualizations
├── data/ # Data files (cleaned CSVs included, raw data excluded)
├── Makefile # Build automation
├── requirements.txt # Python dependencies
└── README.md # This file
Top 30 stations by alert frequency, showing which stations have the most service disruptions
Top 30 stations by average dwell time (crowding proxy), showing passenger volume hotspots
Best Model: Random Forest Classifier
- Accuracy: 70.6% (predicting High/Medium/Low reliability)
- Model: Random Forest with 300 trees, max_depth=6
- Features: 20 selected using mutual information
- Train/Test Split: Temporal (2019-2022 train, 2023-2024 test)
Top 5 Predictive Features:
- Month (0.102) - Strong seasonal patterns
- Snow rolling 3-day average (0.093) - Recent snow patterns matter
- Snow volatility (0.091) - Snow std over 7 days
- Alert frequency (0.083) - Days since last alert
- Alert patterns (0.070) - Monthly alert patterns
SVD Dimensionality Reduction Visualization:
Applied SVD using PCA to reduce the 20-dimensional feature space to 2D
-
Weather alone explains only 3.7% of reliability variance
- Weather is not the main factor of reliability issues.
-
Seasonal patterns
- Month is the #1 predictor
- Winter months (Jan-Feb) have lower reliability
- This makes sense - snow, cold weather, more disruptions
-
Snow patterns
- Recent snow (3-day average) is more predictive than today's snow
- Snow volatility (how much it varies) is also important
-
Alert frequency
- Stations with recent alerts are more likely to have reliability issues
- Pattern learning from alerts helps predict future problems
-
Time-series
- Lag features and rolling averages capture important patterns
- Simple features miss these temporal relationships
-
MBTA Reliability Data (2016-2024)
- Source: MBTA ArcGIS OpenData
- Contains: Daily reliability percentages for Green Line B
- Manual download required
-
Weather Data (2016-2024)
- Source: Visual Crossing API
- Contains: Precipitation, snow, temperature, humidity, etc.
- Downloaded automatically via
make data
-
BU Class Schedule Patterns
- Source: BU standard class meeting times
- Contains: MWF and TR class start times (8am, 10am, 12pm, 2pm, 4pm)
- Encoded in
src/mbta/class_schedules.py
-
MBTA LAMP Performance Data (2019-2024)
- Source: MBTA Performance Data Portal
- Contains: Headway, dwell time, travel time metrics
- Downloaded automatically via
make data
-
MBTA LAMP Alerts Data (2019-2025)
- Source: MBTA LAMP alerts system
- Contains: 4.1M alert records (construction, maintenance, technical problems)
- Downloaded automatically via
make data
- Standardized datetime formats across all datasets
- Handled missing values (filled with zeros where appropriate)
- Filtered for Green Line routes only
- Aggregated to daily level (since reliability data is daily)
- Created temporal features (day of week, month, hour)
- Merged all datasets on date
Created 119 features total, then selected the top 20 using mutual information:
-
Temporal Features
- Day of week, month, hour
- is_weekend, is_monday
- Semester patterns (fall/spring/summer)
-
Weather Features
- Current: precip, snow, snowdepth
- Lagged: snow_lag_1d, snow_lag_7d
- Rolling averages: snow_rolling_3d, snow_rolling_7d
- Volatility: snow_std_7d
-
Class Schedule Features
- Morning class starts (8am-10am)
- Afternoon class starts (12pm-3pm)
- Total class starts per day
-
Alert Features
- Daily counts: construction_alerts, technical_problem_alerts, total_alerts
- Lagged: total_alerts_lag_1d, total_alerts_lag_7d
- Rolling averages: total_alerts_rolling_7d, total_alerts_rolling_14d
- Patterns: alert_pattern_month, days_since_last_alert
-
Interaction Features
- snow_x_classes (snow × class starts)
- alerts_x_classes (alerts × class starts)
Used mutual information to select the top 20 features. This helped reduce overfitting and improve model performance.

