Every time we changed
the machine, and why.
Every football account tells you when its model got something right. Here is the other half: every time we changed the model itself, what was wrong with it, and what we did. No tipster publishes their own corrections. We can't hide ours — the site rebuilds from data twice an hour, and the data does not forget.
Five more leagues get model reads - as young ratings, with no picks
What we found. We re-ran our expected-goals coverage check across every league the collector tracks. Five passed the 80% bar on settled matches - 2. Bundesliga (92%), Primeira Liga (87%), Scottish Premiership (86%), Serie B (83%) and Eredivisie (83%) - meaning their xG data is now reliable enough to build ratings from. MLS (63%), La Liga 2 (73%), Ligue 2 and the European cups stayed below it and remain results-and-stats only.
What we changed. Those five leagues now get model probabilities on the board and fixture pages - appearing automatically once a club has played six games, exactly like everywhere else. Every such number is badged as a young rating: built on a first season of data with no prior seasons behind it. No picks are generated from young leagues, and their accuracy is scored on the calibration page separately from the established six, a split we wrote down before a single one of their predictions had settled.
Why it matters. More of the board with real reads on it - without pretending a first-season rating is as trustworthy as one with years of history. If the young leagues ever earn picks, the public calibration record will be the evidence, not our say-so.
The picks are leaning underdog - we have noticed, and we are testing it
What we found. Since the season opened, the model has priced unfancied sides higher than the market almost across the card - by mid-September nearly every open pick backed the outsider, at big claimed edges. Our own calibration page shows the same thing from another angle: outcomes the model called at 40-70% have been happening more often than it said. In plain English, the model has been under-confident - and an under-confident model sees phantom value in every underdog.
What we changed. Nothing in the model - and that is deliberate. On 12 September, before fitting or testing anything, we wrote down and locked a recalibration exam: it runs in late October, on seasons the fit never touched, with its pass/fail rules fixed in advance. The probabilities change only if it passes. Until it settles, the picks page carries a plain notice whenever the open slate leans this way, so a first-time visitor is told what they are looking at.
Why it matters. Five weeks of early-season data is exactly the sample most likely to mislead: clubs are still inside the warm-up window, and a pattern that loud can be an artefact of that. Correcting the model against it now could bake a bug in permanently. Writing the test down before running it is what keeps us from fooling ourselves - and you.
The board went blank for an afternoon - and why
What we found. For a few hours on 5 September every fixture on the board read 'not enough matches played yet'. The website had been building its ratings from one season at a time - last season until the new one had 200 matches with xG on file, then this season alone. The xG repair below pushed 2026/27 past 200, the site switched seasons, and with clubs on three to five games each, nobody had cleared the six-game warm-up. Meanwhile the picks engine had always carried ratings across seasons with the 20 August early-season blend - so the site and the engine had never quite agreed.
What we changed. The site now builds ratings the same way the picks engine does: every season on file in date order, ratings carried over the summer, and a club's rating pulled toward its prior-season average until it has played six league games this season. Promoted clubs with no prior season start at the league average. The season tables (xPts, goals against xG) still describe this season only.
Why it matters. One set of numbers, not two. The board, fixture pages and team pages now show exactly what the picks are made from, and the early-season caveat already on every fixture page tells you when a rating still leans on last year.
The model was going hungry in two leagues
What we found. Our collector asked for a match's expected goals once, right after the final whistle - but the data provider often posts xG hours or days later. Any match we asked about too early was filed with no xG and never asked about again. By early September only a third of settled Bundesliga games and just over half of Championship games had xG on file. Ratings only learn from games with xG, so in those two leagues the model was reading the season with half the pages missing.
What we changed. The collector now goes back every morning and re-asks for any recent match still missing xG, and we ran a one-off pass over the whole season. All six full-xG leagues are now complete on settled games. Ratings in the Bundesliga and Championship shifted as the missing games landed.
Why it matters. This was not a change to how the model reads a match; it was a fix to what it was allowed to read. It still belongs here, because every published Bundesliga and Championship read before 5 September came from a thinner picture than we thought, and those reads stay on the record exactly as they were made.
Closing odds now taken before kick-off, not after
What we found. Our plausibility watchdog flagged impossible prices on one match (Rennes 2-2 PSG, 23 Aug). For closing odds - the price we grade our picks against - we had been reading the market after full-time, which for that game returned in-play prices near the end of a 2-2 draw. No bet was ever placed on the match and no result was affected, but it was wrong data sitting in the record.
What we changed. Closing odds are now taken from the last price we captured before kick-off - which is what 'closing' is meant to be - instead of a fresh reading after the match. We also removed the corrupted prices for that one fixture.
Why it matters. The closing line is how we honestly check whether our picks beat the market (closing-line value). It has to be the real pre-kick-off price, not an in-play snapshot. This makes that check accurate from here on - and it is the sort of correction we publish rather than bury.
One pick per match (the correlation fix)
What we found. The record logged every conviction selection, so a single match spawned up to sixteen near-identical bets - Asian-handicap lines, team-unders, clean sheet, all the same 'low-scoring' idea in different costumes. Eight days ran to 98 singles across about twelve matches, so one bad afternoon buried the record under correlated red and the whole thing read far worse than the model was actually doing.
What we changed. From now the record keeps one single per match - the model's highest-edge read on that fixture. We also brought the existing record onto the same footing one time: kept the best-edge pick per past match, dropped the correlated duplicates. That moved the settled paper record from about -36 reference units to roughly break-even, which is the model's true per-match story. The full pre-change ledger is kept untouched as a backup.
Why it matters. A record that counts one correlated call sixteen times is not an honest record of the model's calls - it is noise that makes an ordinary, mildly-negative run look like a collapse. This is the picks-correlation lesson (count matches, not chips) applied at the source, and it lines the record up with what the model actually does: one best read per match. We broke the append-only rule once, on purpose and in public, because leaving a misleading record standing would have been the bigger dishonesty.
Early-season rating blend
What we found. For a club's first few games the ratings leaned on last season's final handful of matches - a tiny, distorted sample. It crowned Levante's attack above Barcelona's, and read Barcelona as a coin-flip away at Elche.
What we changed. While a club has played fewer than six league games this season, its rating is now pulled toward its full prior-season average. We tested it across four past seasons before shipping - it improved the model's honesty in every one, and left mature-season games untouched to five decimal places.
Why it matters. It attacks the model's biggest documented weakness at the root, rather than papering over it. A later check (23 Aug) confirmed it costs a sliver of closing-line value in that early window and changes nothing about the no-profit verdict - it earns its place on honesty, not profit.
Distrust guard on lopsided favourites
What we found. Early in the season the model sometimes faded a clear market favourite by a wide margin - rating a side the market made a 71% favourite as barely better than a coin-flip.
What we changed. Added a guard that suppresses the model's own picks on any fixture where it fades a strong market favourite by 15 points or more. The read still shows on the page; it just cannot reach our record.
Why it matters. In exactly those spots we trust the market more than our own model. The guard keeps the artifact out of the record without pretending the model never produced it - honesty cuts both ways.
Asian-handicap and goal-line pricing fix
What we found. Some handicap and total lines were being priced with the wrong push-adjusted maths, inventing edges that were not real - the biggest gap on the board that day was fabricated by the bug.
What we changed. Corrected the maths so every handicap and total line prices the way it actually settles, and added an automated watchdog that fires if a line ladder ever looks impossible again.
Why it matters. A real number in a false frame is the worst kind of error: it passes every test and only shows up when a human looks hard. It reached the board, so it belongs on this page.
How the model works → · What we got wrong this week → · The record →