Case file 04 · Model · Sports analytics
OddsLab
A Premier League match-outcome model built from 24 seasons and 9,120 matches. It's trained offline and runs in the page, so you can price a fixture yourself and compare the answer with what the bookmakers thought.
- Year
- 2026
- Role
- Solo: data, modelling, build
- Read
- 3 min
Context
A match-outcome model for the Premier League, built from a quarter-century of top-flight data: 24 seasons and 9,120 matches. It prices a fixture as home win, draw or away win, with a scoreline distribution underneath.
The point was never to beat the market. I wanted to find out honestly whether a model built from public data and form features could get anywhere near the people who do this for a living, and to show the working.
The problem
Football results are close to a coin weighted three ways, and the bookmakers' prices are a strong baseline that already contains most of what's knowable. A model that can't beat “always pick the home team” is worthless, and a model that seems to beat the market by a lot has almost certainly leaked.
Leakage was the real risk. Form features are calculated from surrounding matches, so it's very easy to let a season's later games inform a prediction about its earlier ones, and end up with a number that looks brilliant and means nothing.
What I built
Twenty form features, a time-ordered split that can't leak, and four models tested against each other instead of one model presented as the answer. Two engines ship. The market engine uses 17 form features plus the bookmakers' implied probabilities with their margin taken out. The no-market engine swaps those odds for Elo ratings, so it can price a fixture that has no odds yet.
The trained model is built into the page and runs in your browser, so it isn't a write-up of a model, it's the model. Pick two teams, nudge the odds, and see where it and a bookmaker disagree. Scorelines come from two Ridge regressions forecasting each side's goals, fed into independent Poisson draws.
Decisions & iterations
The decisions that mattered were all about not fooling myself. The split is by time, not random, and the site has a whole section on leakage: what it is, where it nearly happened here, and what the numbers looked like before I caught it.
The comparison is made the way it should be: against the bookmakers on exactly the same games, not a flattering subset, and next to an “always home” baseline so you can see the floor. Four models are shown side by side instead of one declared winner, because the gaps between them are small enough that a single number would overstate the result.
Outcome
On held-out games the model gets 56.7% right, against the bookmakers' 56.4% on the same fixtures and an “always home” baseline of 43.5%. That's a margin of three tenths of a percentage point. It's the honest headline, and the reason the site talks about limits, not edge. It clearly beats the naive baseline, it lands level with the market, and it says so.