F1-Race-Analysis: The Repo That Turns F1 Telemetry Into Overtake Probabilities
A Streamlit strategy room that fuses live timing, historical priors, and XGBoost predictions to guess when a move is coming before the TV graphics do.
- F1-Race-Analysis treats race strategy as a prediction problem, not a visualization problem.
- Its most interesting move is to turn telemetry into engineered features that approximate racecraft.
- The repo is quietly future-aware, with a 2026 driver lineup that pushes it beyond pure retrospective analysis.
- The codebase feels like an early strategist’s console because it combines live analysis, historical priors, and model inference in one place.
The Pit Wall, Rebuilt in Python
Most F1 tools show you what happened. This repo tries to tell you what is likely to happen next. That shift changes the whole feel of the project: it stops being a dashboard and starts looking like a pit wall console with a prediction habit.
The centerpiece is a Streamlit app that blends live timing, historical features, and a binary classifier to estimate overtakes and future race outcomes. It is a small but sharp idea: if you can encode pace, tyre life, and position pressure well enough, you can make the race a little less mysterious.
Why This Is More Than a Dashboard
| Broadcast analysis | F1-Race-Analysis |
|---|---|
| Shows what already happened | Estimates what is likely to happen next |
| Built for spectators | Built for inference |
| Focuses on charts and replay | Combines live analysis with future predictions |
| Explains the race after the fact | Turns timing data into strategic signals |
That difference matters because the app is organized around use, not archive. The user does not just inspect laps. They move between live analysis and future prediction, which makes the experience feel operational.
In product terms, this is the repo’s thesis. It does not compete on prettier charts. It competes on whether it can help a fan or analyst think more like a strategist.
How the Prediction Stack Is Wired
The architecture is simple to describe and easy to miss in practice. Notebook experiments feed the model work, the processed features land in Parquet, the trained classifier is serialized as a pickle, and the Streamlit app acts as the last layer that stitches everything together.
# Core flow, simplified
laps = load_fastf1_session()
features = engineer_features(laps)
model = joblib.load("f1_overtake_model.pkl")
proba = model.predict_proba(features)[:, 1]
render_dashboard(proba, historical_features.parquet)
That separation is the quiet strength of the repo. Research lives in notebooks. Inference lives in a dedicated module. Presentation lives in the app. For a solo project, that is already a healthier shape than most hobby dashboards.
The Feature Engineering Is the Real Product
The most important function in the codebase is not the model itself. It is engineer_features(laps). That is where raw lap data gets translated into signals the model can actually use.
Rolling pace, tyre advantage, position delta, recent finish form, and DNF history are not glamorous features. They are proxies. That is what makes them useful. The model cannot see grip or intent directly, so the code builds substitutes that approximate pressure, degradation, and momentum.
This is where the project stops looking like a stats page and starts looking like a machine-learning system. The value is not in observing a lap. It is in encoding the race state in a way the model can read.
Why the 2026 Grid Is a Fascinating Design Choice
The 2026 driver lineup is the most revealing design decision in the repo. Hardcoding a future grid is a small act, but it says a lot. The app is not just built for the present season. It is already making room for a season that the live data ecosystem does not fully know yet.
| Current-season thinking | Future-aware thinking |
|---|---|
| Depends on already published entries | Seeds likely entries before the data layer catches up |
| Stays inside retrospective analysis | Extends into speculative forecasting |
| Treats the dataset as fixed | Treats the grid as something to be modeled ahead of time |
| Optimizes for reporting | Optimizes for anticipation |
That choice matters because it nudges the app from analysis toward scenario planning. It is a reminder that predictive tools often have to bootstrap their own future context before the official sources are complete.
The move is also pragmatic. If the dashboard is meant to support future-race predictions, it cannot wait on perfect upstream data. It needs a usable placeholder for tomorrow’s field, even if that means seeding the model with a manually maintained lineup.
What the Model Can and Cannot Know
The classifier is powerful, but it is still a classifier. With an XGBoost model, binary logistic output, 300 trees, and a max depth of 6, the repo is making a serious prediction engine. It is not making a certainty machine.
XGBClassifier(
objective="binary:logistic",
n_estimators=300,
max_depth=6
)
That distinction is important. The model can learn patterns from historical race behavior, but it cannot invent truth the data never contained. If the features are noisy, biased, or incomplete, the probability score will still look confident.
So the real question is not whether the model is clever. It is whether the feature stack captures enough of racecraft to make its probabilities meaningful. In this repo, that is the right problem to solve.
How It Compares to the F1 Data Ecosystem
| Project | Role in the stack | What it gives you |
|---|---|---|
| FastF1 | Data access and analysis library | Telemetry, timing, and pandas-friendly race data |
| OpenF1 | Live and historical API | Real-time telemetry and race data without authentication |
| jolpica-f1 | Historical data API | Ergast-compatible historical records |
| F1-Race-Analysis | Application layer | Feature engineering, prediction, and a live strategy UI |
That comparison is the key to understanding the repo’s position. FastF1 and OpenF1 are enablers. This project is an opinionated application that sits on top of them and tries to operationalize the data.
It is not trying to be the broadest source of F1 information. It is trying to be the smartest layer between raw feeds and strategic judgment. That is a smaller scope, but a sharper one.
The Prototype Ceiling
- The codebase is functional, but it is still monolithic in places.
- The project has no visible tests or CI, so reliability depends heavily on manual discipline.
- The absence of a formal license and wider community footprint keeps it in prototype territory.
- That does not weaken the idea, because the ambition is already ahead of the polish.
That tension is actually part of the appeal. You can see the shape of a better product inside the rough edges. The repo has enough structure to feel real, but enough looseness to feel like it is still discovering itself.