Skip to content

betting_combat.rounds.features.clock

Group 5: model outputs as features, fitted walk-forward (a row in year Y only ever sees models fitted on fights before 1 January of Y). Known pre-fight / live from the round number.

Port of the research ufc/dataset/models.py (ModelFeatures); the numbers must stay identical.

m_p_dec_prefight P(goes to the judges) before the fight, from fighter history alone: a logistic regression on both fighters’ decision / finish / durability rates, 5 rounds and women’s division, refitted every year. m_clock_S S(m | x): share of this fight’s possible finishes that would still come after minute m = 5 x rounds fought. The Bayesian piecewise-exponential finish-time curve (strategy 2’s C3 curve): minute-by-minute baseline hazard with a random-walk prior, partially pooled division effects, covariate multipliers; refitted every year. m_p_dec_now P(decision | no finish by minute m) = p0 / (p0 + (1 - p0) S(m)) — strategy 2’s fair value, with p0 the bookmakers’ P(decision) where it exists, else m_p_dec_prefight. m_p_finish_round P(finish in the round about to start | no finish yet) = (1 - p0) (S(m) - S(m + 5)) / (p0 + (1 - p0) S(m)). m_p0, m_S_1..m_S_4 the fight’s p0 and its whole curve (S at the end of rounds 1-4), so any future round can be priced (used by the joint model; not model inputs).

The curve. The sampler is strategy 2’s own (models.training.fit.fit_finish_curve: the same NumPyro model, design and seed handling as the research BayesPE; its draws equal the research’s bit for bit), run with every draw kept (thin=1), as the research did. S is evaluated here, not with models.distance.FinishCurve: the research averaged S over the draws in the sampler’s own precision (float32 draws, products such as steps x sigma_rw taken in float32 before the float64 hazard), while FinishCurve converts the draws to float64 first — up to 4e-10 apart in S on the 2012 curve. ClockCurve.s_cond_samples is the research’s vectorised formula, operation for operation.

The research fitted every year (~15-60 s each) and cached the result; production takes the cache as an argument (dataset.RoundDataset.build(model_features=...)).

Classes

ClockCurve

A fitted finish-time curve: the training design stats and every posterior draw.

fit

fit(fin: pd.DataFrame, samples: int, warmup: int, seed: int) -> ClockCurve

NUTS on the finishes fin (columns COVARIATES, division, bin, K), every draw kept.

s_cond

s_cond(rows: pd.DataFrame) -> np.ndarray

S(m | x) averaged over the posterior draws.

s_cond_samples

s_cond_samples(rows: pd.DataFrame) -> np.ndarray

(draws, rows) matrix of S(m | x) across posterior draws (rows: m, K, division, COVARIATES).

ClockFeatures

build

build(rows: pd.DataFrame, book_p_dec: pd.Series) -> pd.DataFrame

rows: fight_id, round_idx, event_date. book_p_dec: per fight_id (NaN where none).

fight_table

fight_table() -> pd.DataFrame

Every 3/5-round bout since 2001 with its fight-level covariates (as-of profiles).

fit_curve

fit_curve(fin: pd.DataFrame, seed: int, tries: int = 4, tol: float = 0.08) -> ClockCurve

Fit the Bayesian curve and check it converged: on its own training finishes, the mean fitted S(5) (share of finishes after round 1) must match the actual share within tol, per format. A stuck chain fails this badly; it is refitted with a new seed and a longer warm-up.

fit_year

fit_year(f: pd.DataFrame, year: int) -> ClockYear

Year year’s models, fitted on the bouts of f (fight_table) before 1 January of that year: the p0 logistic model and the Bayesian curve (seed year).

year_frame

year_frame(fitted: ClockYear, rows: pd.DataFrame, f: pd.DataFrame, book_p_dec: pd.Series) -> pd.DataFrame

Group 5 for rows (fight_id, round_idx; all in fitted.year) from that year’s fitted models. f: fight_table (holding the rows’ bouts); book_p_dec per fight_id (NaN where none).

ClockYear

One year’s fitted Group 5 models (ClockFeatures.fit_year): everything ClockFeatures.year_frame needs to price that year’s rows, upcoming bouts included.