A probability product, not a betting product

The model grades itself

Every probability the odds model would have produced, scored against what actually happened. Out of sample, walk-forward, with a temporal split. When the model is wrong, this page shows it. Keeping score in public is the point.

Held-out calibration

Predicted vs actual

Each point is a probability decile on the held-out races. On the dashed line is perfect; below it is overconfident.

00252550507575100100perfect calibrationpredicted win probability (%)actual win frequency (%)

The honest read

Well-calibrated low, overconfident high

Below 50% predicted, the model is well-calibrated: when it says 24%, horses win 26%. Above 50% it is overconfident: when it says 84%, horses win 58%. Strong prior records overstate certainty in a noisy sport. Narrowing that high-end gap is v2's job, and it is stated here rather than hidden.

Scores on held-out races

Brier score
0.107
vs 0.130 uniform
Brier skill
17%
better than guessing
Log loss
0.352
lower is better
Held-out entries
84,141
13,592 races

The split, stated

How this avoids grading itself on its own training data

Races run in time order. The model fits its one parameter on the earlier 31,714 races (before race #33,774) and is scored only on the later 13,592 held-out races. Every prediction uses a horse's record from races strictly before the one being predicted, so the outcome is never an input. Method: temporal walk-forward.

Scope

Win-rate-driven probability, validated out of sample. ELO and stat reveals are current-only in our data and would leak past outcomes, so they are excluded from the historical curve; the live odds endpoint adds them and labels them uncalibrated.

Model odds-v1-winrate-core. Backtest computed 1h ago, served precomputed by /api/v1/calibration. Held-out field baseline win rate 16.2%.