Methodology
How Paddock knows what it knows
One engine produces every number on this site. This page is the audit trail: what we measure, what we will not pretend to measure, and how every threshold and model was checked against real outcomes rather than asserted.
First, what we will not do
The honest-data principle
A Gigling's four stats are hidden, shown only as min-max ranges that narrow as the horse races. Paddock never collapses an unrevealed range into a single number. A horse showing 50 to 100 has an unknown stat; its midpoint is not information. Every value on the site is one of three things: a true revealed value, an explicit range with its reveal percentage, or an estimate clearly labeled with its method. When something cannot be known, the interface says so.
A full-population read
The study behind the scores
The scoring weights come from a study of every resolved race at the time: 4,537 races, 30,288 entries. The weights are frozen from that snapshot, so the constants stay fixed even as live data grows (now past 45,312 resolved races); 4,537 is the studied population, not a stale count. The strongest single signal is raw stat quality: race winners average 3.8% above the field on all four stats. Among traits, Surger is the alpha: a 23.19% win rate when active versus the study's 14.18% baseline, a 1.63x lift overall. Several traits change sign by distance, which is why track fit is computed per length rather than globally: Surger is 1.63x across all tracks but 1.69x at 1200m, and Closer's edge appears only at 2400m and longer.
The two numbers
Confirmed quality, upside, and shrinkage
Confirmed quality uses only revealed information: revealed stat values, revealed trait star levels weighted by their study lift, and an actual win rate. Upside is the opposite: for unrevealed horses it reads rarity, the traits a horse carries from birth, and races remaining, and it is labeled potential, never proof. Win rate everywhere is Bayesian-shrunk toward the study's 14.18% baseline, the same rate the study measured, so a 2-for-3 horse does not outrank a 20-for-60 horse on three lucky races. The raw record is always shown beside the shrunk number.
The part that is hard to fake
Validated out of sample, everywhere
A threshold that is merely rare is not the same as a threshold that is predictive. So both the scanner's shark flag and the odds model were checked the same disciplined way: walk the races in time order, judge each horse only on its record from races strictly before the one being scored, and measure what actually happened next. The outcome is never an input to the prediction.
Scanner shark flag, out of sample
| Line | Win rate | Lift |
|---|---|---|
| 0.25 | 45.7% | 3.06x |
| 0.28 | 48.6% | 3.26x |
| 0.30 ★ | 50.9% | 3.41x |
| 0.33 | 53.5% | 3.59x |
| 0.35 | 55.3% | 3.71x |
Flagged horses (at 0.30) win 50.9% of their next races versus the full-entry 14.9% baseline (the win rate across all finished entries, distinct from the study's 14.18%). Predictive, not cosmetic; the smooth gradient is the signature of real signal.
Odds model, held-out calibration
- Brier score
- 0.107 vs 0.130 uniform
- Brier skill
- 17% better than guessing
- Split
- 31413 train / 13464 held out
Well-calibrated below 50%, overconfident above. Shown, not hidden, on the calibration page.
These two findings are one story. The odds model is overconfident on its favorites: when it says 84%, those horses win about 58%. Its high-confidence picks, averaging a 69% predicted chance, actually win 52%. That is the same number the shark cohort wins, near 50%. Whether you define a strong horse relatively, as the model's favorite, or absolutely, as a shark, the truth is identical: elite Giglings are strong but beatable, and a naive favorite-take overstates them. The scanner and the odds model show the same reality from two angles.
Real, but modest, and we say so
Distance fit, validated the same way
Distance fit used to sit in the "what we know" list on assertion alone. It now earns the same out-of-sample table as the shark flag and the odds model, over 8,398 resolved races and 55,683 entries, reproducible from scripts/study-distance-fit.mts. (That 8,398 is a later frozen snapshot than the scoring study's 4,537 above: both are audit snapshots taken when each study was run, not live counts, which is why the current total, now past 45,312, is higher than either.) Fit is outcome-independent (it reads only revealed stats and traits, never wins, ELO, or finish), so unlike ELO it cannot leak the result; the relationship even strengthens as a horse's stats reveal more, which is the signature of a real effect rather than an artifact.
Raw gradient: higher fit, better finish
| Fit at the track | Mean finish pctile |
|---|---|
| Lowest fit decile (~24) | 0.68 |
| Middle decile (~50) | 0.53 |
| Highest fit decile (~73) | 0.38 |
Monotonic across all ten deciles, Pearson r is -0.29 over the full set and -0.33 among the most-revealed horses. But most of this is horse quality: good horses carry high fit at every distance and finish well everywhere, so this raw link overstates fit's own contribution.
Within-horse, quality-controlled penalty
| Fit pts below own best | Finish penalty |
|---|---|
| At or near best (< 3) | -0.007 |
| 5 to 10 below | +0.004 |
| 15 to 20 below ★ | +0.009 |
| 20+ below | +0.009 |
Each race scored against the same horse's mean across its other races (leave-one-out), so horse quality is removed and a race never baselines on its own outcome. The honest, isolated effect of distance fit: real and monotonic, but small, roughly a tenth of a finishing position end to end. It firms up (+0.009) at 15 points below best, which is exactly where the scanner's weak-fit caution begins.
So the scanner treats fit as a heads-up, not a stop sign. Below 3 points off a horse's own best is within noise and says nothing. From 3 to 15 points is a quiet "off best" note. Only at 15 points or more below best, and only when that gap is at least half the horse's own fit spread, does it raise a soft caution. The thresholds are these numbers, not hand-picked examples.
Comps or silence
Valuation, and when we stay quiet
A valuation band is the interquartile range of comparable sales: same rarity, similar confirmed quality, similar reveal state. Below 3 comparable sales we show no band at all and say the comps are thin. Between 3 and 4 comps we show the band but flag it low-confidence. We never manufacture a precise number from a market that is too quiet to support one.
The honest gap
What cannot be known
Paddock is the source of truth for what is knowable and persistent about a horse: its revealed stats and ranges, traits and tiers, study-measured lifts, confirmed quality and upside, distance fit, and its win rate and ELO from the actual on-chain record. Three things, by contrast, no pre-race tool can see. We name them rather than pretend otherwise.
First, daily readiness. A horse has a daily race cap and becomes exhausted, but the public cooldown field reads 0 even for an exhausted horse, so Paddock cannot claim live readiness and does not.
Second, point-in-time history. Our data holds current ELO and current stat reveals, not their values on past dates, which is exactly why the odds backtest excludes them. Including current values to grade past races would leak the outcome. The scanner states this limit on every verdict.
Third, and most significant: items and sabotage. Players can spend consumables mid-race: butterflies that add 5 to 10 percent speed to their own horse, dung that takes 5 to 10 percent off a rival, each lasting five seconds, with faction-matched versions hitting harder. Only one item applies per play and it fires at the next resolve interval, so the effect is a series of timed nudges, not a permanent multiplier. But these are decided in real time by other players during the race, are not recorded on-chain, and leave no trace Paddock can read, before or after. Paddock does not model them. The scores describe the race that horse quality and distance fit would produce; item play is a layer of live human agency on top of that, invisible to any pre-race analysis. This is also why a horse's win rate is noisier than pure ability: its record already silently includes races where items swung the result.
Safety
Your assets never move
Paddock's read surfaces never touch a wallet. The optional auto-racer signs exactly one kind of transaction, a zero-value free-race entry, and the racing contract only reads ownership: it never transfers a Gigling and never needs an approval. This is proven on-chain, not asserted, with a committed analysis, a re-runnable forensics harness, and a signer-rejection test that holds the safety guard to account. The full write-up is in SECURITY.md in the repository.
Every figure here is reproducible from the public API and the scripts in the repository. The model grades itself in public on the calibration page.