We publish whether our
probabilities were any good.
The record page tracks the calls. This page tracks every probability the model publishes — scored after the fact, in public, wins and misses alike.
What calibration means, plainly: when the model says something is 60% likely, it should happen roughly 60% of the time — not 45%, not 75%. A tipster can hide behind a hot streak; a calibration table cannot. It is the one number on this site that tells you whether the percentages mean anything at all, which is why we score it in public and link it from every miss on what we got wrong.
How this page cannot cheat
Every probability here was written to an append-only log before kick-off, from the ratings as they stood at that moment. Nothing is recomputed after results are known — recomputing would smuggle the answer into the question. Where several pre-match snapshots exist, the last one before kick-off is the one judged. The markets measured (match result, over 2.5 goals, both teams to score) and the 50 settled-prediction threshold were fixed before any data existed and do not move mid-season.
When the model says X%, does X% happen?
1010 settled predictions · 230 awaiting results · overall Brier score 0.2174 (lower is better; 0.25 is what coin-flipping 50% at everything scores) · log-loss 0.6247 (0.693 is that same coin-flip — log-loss just punishes confident misses harder).
Each dot is a probability band; bigger dots hold more predictions. The dashed line is perfect honesty — a dot on it means when the model said that, it happened that often. Above the line the model was too cautious; below it, too confident.
| Model said | Predictions | Avg said | Actually happened | Gap |
|---|---|---|---|---|
| 10–20% | 16 | 18.0% | 12.5% | -5.5 pts |
| 20–30% | 281 | 26.7% | 23.5% | -3.2 pts |
| 30–40% | 166 | 34.2% | 30.7% | -3.5 pts |
| 40–50% | 234 | 45.4% | 50.4% | +5.1 pts |
| 50–60% | 247 | 54.4% | 61.5% | +7.1 pts |
| 60–70% | 63 | 62.9% | 73.0% | +10.1 pts |
| 70–80% | 2 | 75.3% | 100.0% | +24.7 pts |
| 80–90% | 1 | 81.0% | 100.0% | +19.0 pts |
| Market | n | Brier | Log-loss |
|---|---|---|---|
| Home win | 202 | 0.2249 | 0.6411 |
| Draw | 202 | 0.1794 | 0.5444 |
| Away win | 202 | 0.2084 | 0.6045 |
| Over 2.5 goals | 202 | 0.2328 | 0.6578 |
| Both teams to score | 202 | 0.2414 | 0.6758 |
Against the market's de-vigged closing price, over the 1003 settled prediction(s) where a full close was captured: our log-loss 0.6244 versus the market's 0.6052 (lower is better). The closing price scored better here — the market is a hard baseline, and we do not pretend to beat it.
The young leagues, scored separately
Five leagues joined the model on 13 September 2026 — 2. Bundesliga, Serie B, Eredivisie, Primeira Liga and the Scottish Premiership — because their expected-goals data proved reliable. Their ratings are built on a first season of data only, so their probabilities are badged young wherever they appear, no picks are generated from them, and they are scored here separately: the numbers above are the established six leagues alone, and stay that way.
10 settled prediction(s) · 35 awaiting results · Brier 0.2529 · log-loss 0.7038 (0.25 and 0.693 are the coin-flip baselines; lower is better). Small samples early — judge these numbers only as their n grows.