Why Logistic Regression outperformed eight other models in my F1 research
How LOCO evaluation revealed what deep learning got wrong.
The Setup
For six months, I co-authored a research manuscript benchmarking nine machine learning and deep learning models on a single problem: predicting active aerodynamic system deployment across Formula 1 circuits. The dataset was 159,712 telemetry frames from Monaco, Monza, Silverstone, and Suzuka.
The evaluation protocol was Leave-One-Circuit-Out (LOCO): train on three circuits, test on the fourth. This is a stricter test than standard random split cross-validation. A model can't memorise a track's characteristics — it must actually generalise.
The models ranged from Logistic Regression to LSTM networks. We expected the deep learning architectures to win. They didn't.
The Result That Surprised Everyone
Logistic Regression achieved AUC 0.963 ± 0.021. It was the best-performing model across all four circuits.
The LSTM, which we expected to dominate on sequential telemetry data, struggled. The standard MLP was inconsistent. Random Forest and Gradient Boosting were competitive but slightly below LR on average.
This isn't a fluky result. It held across every circuit and every fold.
Why Did This Happen?
Speed memorisation in recurrent architectures.
When an LSTM trains on Monaco, Monza, and Silverstone, it learns what "Formula 1 speed at turn X" looks like for those circuits. When evaluated on Suzuka — a fundamentally different layout — the model applies circuit-specific priors to a new context it hasn't seen. The sequential memory that makes LSTMs powerful for some time-series problems becomes a liability here.
Logistic Regression doesn't build this kind of memory. It learns feature weights — relationships between telemetry variables and the target label — and those relationships are more circuit-agnostic than the raw sequential patterns.
This is the core finding: when distribution shift is the primary challenge, model simplicity can be an advantage. Feature-level generalisation outperforms sequential pattern memorisation when the sequence context is domain-specific.
What Integrated Gradients Told Us
We used Integrated Gradients interpretability to understand which features each model was relying on. The LR model's feature weights aligned with physically intuitive variables: speed, throttle position, and DRS state indicators. These are the actual mechanical conditions under which active aero systems deploy.
The LSTM attention patterns were harder to interpret and showed circuit-specific signatures — evidence of memorisation.
This combination of LOCO evaluation and interpretability analysis is what transformed a benchmark paper into a diagnostic paper.
What This Means for ML in Domain-Shifted Settings
The lesson extends beyond F1. In any domain where test conditions differ systematically from training conditions:
- Do not assume more complexity equals better generalisation
- Evaluate feature attributions to check for distribution-specific shortcuts
- Design evaluation protocols that force the model to face the distribution shift explicitly (LOCO, leave-one-domain-out, etc.)
The manuscript is currently under review. I'll link the paper here when it's published.
All benchmarking was done in Python using Scikit-Learn, TensorFlow, and Captum for interpretability. Data sourced from the FastF1 telemetry API.