betting_combat.rounds.features.clock
Group 5: model outputs as features, fitted walk-forward (a row in year Y only ever sees models fitted on fights before 1 January of Y). Known pre-fight / live from the round number.
Port of the research ufc/dataset/models.py (ModelFeatures); the numbers must stay
identical.
m_p_dec_prefight P(goes to the judges) before the fight, from fighter history alone: a logistic regression on both fighters’ decision / finish / durability rates, 5 rounds and women’s division, refitted every year. m_clock_S S(m | x): share of this fight’s possible finishes that would still come after minute m = 5 x rounds fought. The Bayesian piecewise-exponential finish-time curve (strategy 2’s C3 curve): minute-by-minute baseline hazard with a random-walk prior, partially pooled division effects, covariate multipliers; refitted every year. m_p_dec_now P(decision | no finish by minute m) = p0 / (p0 + (1 - p0) S(m)) — strategy 2’s fair value, with p0 the bookmakers’ P(decision) where it exists, else m_p_dec_prefight. m_p_finish_round P(finish in the round about to start | no finish yet) = (1 - p0) (S(m) - S(m + 5)) / (p0 + (1 - p0) S(m)). m_p0, m_S_1..m_S_4 the fight’s p0 and its whole curve (S at the end of rounds 1-4), so any future round can be priced (used by the joint model; not model inputs).
The curve. The sampler is strategy 2’s own (models.training.fit.fit_finish_curve: the
same NumPyro model, design and seed handling as the research BayesPE; its draws equal
the research’s bit for bit), run with every draw kept (thin=1), as the research did.
S is evaluated here, not with models.distance.FinishCurve: the research averaged S over
the draws in the sampler’s own precision (float32 draws, products such as steps x sigma_rw
taken in float32 before the float64 hazard), while FinishCurve converts the draws to
float64 first — up to 4e-10 apart in S on the 2012 curve. ClockCurve.s_cond_samples is
the research’s vectorised formula, operation for operation.
The research fitted every year (~15-60 s each) and cached the result; production takes the
cache as an argument (dataset.RoundDataset.build(model_features=...)).
Classes
ClockCurve
A fitted finish-time curve: the training design stats and every posterior draw.
fit
fit(fin: pd.DataFrame, samples: int, warmup: int, seed: int) -> ClockCurveNUTS on the finishes fin (columns COVARIATES, division, bin, K), every draw kept.
s_cond
s_cond(rows: pd.DataFrame) -> np.ndarrayS(m | x) averaged over the posterior draws.
s_cond_samples
s_cond_samples(rows: pd.DataFrame) -> np.ndarray(draws, rows) matrix of S(m | x) across posterior draws (rows: m, K, division, COVARIATES).
ClockFeatures
build
build(rows: pd.DataFrame, book_p_dec: pd.Series) -> pd.DataFramerows: fight_id, round_idx, event_date. book_p_dec: per fight_id (NaN where none).
fight_table
fight_table() -> pd.DataFrameEvery 3/5-round bout since 2001 with its fight-level covariates (as-of profiles).
fit_curve
fit_curve(fin: pd.DataFrame, seed: int, tries: int = 4, tol: float = 0.08) -> ClockCurveFit the Bayesian curve and check it converged: on its own training finishes, the
mean fitted S(5) (share of finishes after round 1) must match the actual share within
tol, per format. A stuck chain fails this badly; it is refitted with a new seed
and a longer warm-up.
fit_year
fit_year(f: pd.DataFrame, year: int) -> ClockYearYear year’s models, fitted on the bouts of f (fight_table) before
1 January of that year: the p0 logistic model and the Bayesian curve (seed year).
year_frame
year_frame(fitted: ClockYear, rows: pd.DataFrame, f: pd.DataFrame, book_p_dec: pd.Series) -> pd.DataFrameGroup 5 for rows (fight_id, round_idx; all in fitted.year) from that year’s
fitted models. f: fight_table (holding the rows’ bouts); book_p_dec per
fight_id (NaN where none).
ClockYear
One year’s fitted Group 5 models (ClockFeatures.fit_year): everything
ClockFeatures.year_frame needs to price that year’s rows, upcoming bouts included.