F1-Race-Analysis: The Repo That Turns F1 Telemetry Into Overtake Probabilities

A Streamlit strategy room that fuses live timing, historical priors, and XGBoost predictions to guess when a move is coming before the TV graphics do.

8 min read • View on GitHub • More from Daksh3077

A wide strategist’s desk with a laptop showing lap times, a second screen showing a rising probability curve, and paper race maps marked with annotations. It explains the core idea that this repository treats race analysis as a decision engine, not just a dashboard.
The project’s real ambition is not to display timing data. It is to turn that data into a strategist’s signal.
Key Takeaways

The Pit Wall, Rebuilt in Python

Most F1 tools show you what happened. This repo tries to tell you what is likely to happen next. That shift changes the whole feel of the project: it stops being a dashboard and starts looking like a pit wall console with a prediction habit.

The centerpiece is a Streamlit app that blends live timing, historical features, and a binary classifier to estimate overtakes and future race outcomes. It is a small but sharp idea: if you can encode pace, tyre life, and position pressure well enough, you can make the race a little less mysterious.

Why This Is More Than a Dashboard

Broadcast analysisF1-Race-Analysis
Shows what already happenedEstimates what is likely to happen next
Built for spectatorsBuilt for inference
Focuses on charts and replayCombines live analysis with future predictions
Explains the race after the factTurns timing data into strategic signals

That difference matters because the app is organized around use, not archive. The user does not just inspect laps. They move between live analysis and future prediction, which makes the experience feel operational.

In product terms, this is the repo’s thesis. It does not compete on prettier charts. It competes on whether it can help a fan or analyst think more like a strategist.


How the Prediction Stack Is Wired

The system is a pipeline, not a monolith. Raw telemetry becomes features, features become predictions, and predictions become the dashboard.

The architecture is simple to describe and easy to miss in practice. Notebook experiments feed the model work, the processed features land in Parquet, the trained classifier is serialized as a pickle, and the Streamlit app acts as the last layer that stitches everything together.

# Core flow, simplified
laps = load_fastf1_session()
features = engineer_features(laps)
model = joblib.load("f1_overtake_model.pkl")
proba = model.predict_proba(features)[:, 1]
render_dashboard(proba, historical_features.parquet)

That separation is the quiet strength of the repo. Research lives in notebooks. Inference lives in a dedicated module. Presentation lives in the app. For a solo project, that is already a healthier shape than most hobby dashboards.

The Feature Engineering Is the Real Product

The most important function in the codebase is not the model itself. It is engineer_features(laps). That is where raw lap data gets translated into signals the model can actually use.

A close-up of a single telemetry line splitting into layered branches for pace, tyre advantage, and position trend, then feeding a compact model core that emits an overtake probability signal. It explains how the repo converts raw race data into machine-learning features.
The real trick is not raw telemetry. It is the translation from telemetry into features that resemble racecraft.

Rolling pace, tyre advantage, position delta, recent finish form, and DNF history are not glamorous features. They are proxies. That is what makes them useful. The model cannot see grip or intent directly, so the code builds substitutes that approximate pressure, degradation, and momentum.

This is where the project stops looking like a stats page and starts looking like a machine-learning system. The value is not in observing a lap. It is in encoding the race state in a way the model can read.

Why the 2026 Grid Is a Fascinating Design Choice

The 2026 driver lineup is the most revealing design decision in the repo. Hardcoding a future grid is a small act, but it says a lot. The app is not just built for the present season. It is already making room for a season that the live data ecosystem does not fully know yet.

Current-season thinkingFuture-aware thinking
Depends on already published entriesSeeds likely entries before the data layer catches up
Stays inside retrospective analysisExtends into speculative forecasting
Treats the dataset as fixedTreats the grid as something to be modeled ahead of time
Optimizes for reportingOptimizes for anticipation

That choice matters because it nudges the app from analysis toward scenario planning. It is a reminder that predictive tools often have to bootstrap their own future context before the official sources are complete.

The move is also pragmatic. If the dashboard is meant to support future-race predictions, it cannot wait on perfect upstream data. It needs a usable placeholder for tomorrow’s field, even if that means seeding the model with a manually maintained lineup.

What the Model Can and Cannot Know

The classifier is powerful, but it is still a classifier. With an XGBoost model, binary logistic output, 300 trees, and a max depth of 6, the repo is making a serious prediction engine. It is not making a certainty machine.

XGBClassifier(
  objective="binary:logistic",
  n_estimators=300,
  max_depth=6
)

That distinction is important. The model can learn patterns from historical race behavior, but it cannot invent truth the data never contained. If the features are noisy, biased, or incomplete, the probability score will still look confident.

So the real question is not whether the model is clever. It is whether the feature stack captures enough of racecraft to make its probabilities meaningful. In this repo, that is the right problem to solve.

How It Compares to the F1 Data Ecosystem

ProjectRole in the stackWhat it gives you
FastF1Data access and analysis libraryTelemetry, timing, and pandas-friendly race data
OpenF1Live and historical APIReal-time telemetry and race data without authentication
jolpica-f1Historical data APIErgast-compatible historical records
F1-Race-AnalysisApplication layerFeature engineering, prediction, and a live strategy UI

That comparison is the key to understanding the repo’s position. FastF1 and OpenF1 are enablers. This project is an opinionated application that sits on top of them and tries to operationalize the data.

It is not trying to be the broadest source of F1 information. It is trying to be the smartest layer between raw feeds and strategic judgment. That is a smaller scope, but a sharper one.

The Prototype Ceiling

That tension is actually part of the appeal. You can see the shape of a better product inside the rough edges. The repo has enough structure to feel real, but enough looseness to feel like it is still discovering itself.