2026-10-05 · Mystique Sports 玄機波馬 研究團隊

Ratings, Speed Figures and Market Odds: A Map for Racing Data Science

Three languages describe the same race. Learning to translate between them is the whole game.

Three languages, one race

Every serious approach to horse racing analysis speaks one of three dialects:

  1. Ratings (Elo-style): each horse carries a single number; beating a rival moves both numbers in opposite directions
  2. Speed figures: each past performance is converted into one time-based number, adjusted for track, distance and condition
  3. Market probability: the odds, cleaned of the bookmaker's margin, read as the crowd's probability estimate

They describe the same reality with different strengths. The craft is in the translation.

Language 1: Elo-style ratings

Borrowed from chess, the idea is simple:

  • Every horse starts at a baseline rating
  • After each race, ratings update based on who beat whom, scaled by margin and strength of opposition
  • A rating difference maps to an expected win probability via a logistic curve

Strengths: self-correcting, handles fields of different sizes naturally, needs no track-specific normalization. Weaknesses: slow to react to rapidly improving young horses; says nothing about why (distance, going, draw).

A practical refinement is decay: weight recent form more heavily (our own system uses EMA-style adaptation — see the EMA article), and context flags: only update fully when the horse ran under comparable conditions.

Language 2: speed figures

Speed figures answer: "How fast did this horse actually run, once we strip out the context?"

The pipeline:

  1. Raw finishing time for the race
  2. Normalize against the course and distance (a 1200m time at Sha Tin means nothing next to a 1200m time at Happy Valley)
  3. Adjust for going — a good-to-firm day makes every raw time look better; the standard fix is a daily track variant estimated from all races that day
  4. Adjust for weight carried, on a researched (not guessed) pounds-per-length scale
  5. Optionally adjust for trouble in running (wide trips, blocked runs) — powerful but subjective

Strengths: directly comparable across horses, courses and seasons. Weaknesses: the adjustments ARE the model. Garbage variant estimates in, garbage figures out. This is why honest pipelines publish their variant methodology.

Language 3: market-implied probability

Decimal odds (O) imply probability (1/O). Summed across a field, this exceeds 1 — the excess is the overround (the operator's margin). To read the market honestly, remove it.

Two standard methods:

  • Proportional (basic): divide each implied probability by the total. Simple, but known to distort — longshots carry disproportionate margin (the favourite–longshot bias)
  • Power method: find an exponent (k) such that (\sum (1/O_i)^k = 1). Better behaved across odds ranges and easy to compute

What remains is the crowd's probability estimate — the strongest single predictor available, and the benchmark any model must be compared against.

The synthesis: calibration and divergence

A model that outputs "Horse A will win" is entertainment. A model that outputs "Horse A wins 22% of the time" is testable.

The workflow we follow at Mystique:

  1. Build probabilities from ratings + speed figures + conditions (gradient boosting on tabular features works well)
  2. Calibrate them (Platt scaling / isotonic regression) on a strictly future validation window, verified with Brier score and calibration curves
  3. Compare against market probabilities from step 3 above: the divergence between model and market is where research questions live
  4. Audit over time: does the model's divergence at race time predict anything about how odds move later? This is the same idea as closing line value — a process metric, not a profit promise

What this map is for

None of these three languages is "the secret". The market is extremely good; beating it consistently is genuinely hard, and anyone promising otherwise is selling something. The point of building your own ratings and figures is understanding: knowing when a horse's price reflects its form, when it reflects a narrative, and being able to say why in numbers.

That is the standard of every article in this series: method public, results reproducible.


This article is for research and education only and is not betting advice. Horse racing involves risk. 18+ (or the legal age in your jurisdiction).

感謝閱讀。如對研究方法有疑問,歡迎透過會員中心的「聯絡研究團隊」與我們交流。

← 回到研究隨筆列表