A probability product, not a betting product
The model grades itself
Every probability the odds model would have produced, scored against what actually happened. Out of sample, walk-forward, with a temporal split. When the model is wrong, this page shows it. Keeping score in public is the point.
Held-out calibration
Predicted vs actual
Each point is a probability decile on the held-out races. On the dashed line is perfect; below it is overconfident.
The honest read
Well-calibrated low, overconfident high
Below 50% predicted, the model is well-calibrated: when it says 24%, horses win 26%. Above 50% it is overconfident: when it says 84%, horses win 58%. Strong prior records overstate certainty in a noisy sport. Narrowing that high-end gap is v2's job, and it is stated here rather than hidden.
Scores on held-out races
- Brier score
- 0.107
- vs 0.130 uniform
- Brier skill
- 17%
- better than guessing
- Log loss
- 0.352
- lower is better
- Held-out entries
- 84,141
- 13,592 races
The split, stated
How this avoids grading itself on its own training data
Races run in time order. The model fits its one parameter on the earlier 31,714 races (before race #33,774) and is scored only on the later 13,592 held-out races. Every prediction uses a horse's record from races strictly before the one being predicted, so the outcome is never an input. Method: temporal walk-forward.
Scope
Win-rate-driven probability, validated out of sample. ELO and stat reveals are current-only in our data and would leak past outcomes, so they are excluded from the historical curve; the live odds endpoint adds them and labels them uncalibrated.
Model odds-v1-winrate-core. Backtest computed 1h ago, served precomputed by /api/v1/calibration. Held-out field baseline win rate 16.2%.