Skip to content

betting_combat.rounds.features.dataset

RoundDataset: assemble the three tables and their column catalogue.

Port of the research ufc/dataset/build.py; the numbers must stay identical. The research wrote files; production returns the tables and the caller stores them.

dataset1 one row per (fight, round start); keys + context + Group 1 (both
fighters) + Group 2 (matchup) + Group 4 (pre-fight odds) + targets +
Group 5 (clock model) + Group 6 (live markets, when given)
dataset2 exactly dataset1's rows and columns + Group 3 (rounds fought so far)
fighters one row per (fighter, fight): that fighter's Group 1 profile alone and
fighter-level targets (win this fight, win this and the next, finish,
get finished)
catalogue every column: its group, and when it is known (pre-fight / live at the
round break from Kalshi / live only with a stats feed / outcome)
model_features Group 5 as used (the research's cache, extended where rows were missing)

Rows kept: fights from 2008 on, won or lost (draws, no contests, DQs, overturned and ‘other’ results removed), 3- or 5-round fights of 5-minute rounds, with round stats for every round that started. Fighter histories use every fight back to 1994.

A/B: fighter A is UFCStats slot 1 or 2 by a fixed hash of the fight id (UFCStats lists the winner first 63% of the time, so the slots must not be used as they come).

Two inputs come from elsewhere: model_features the Group 5 cache (clock.ClockFeatures output). Rows it lacks are fitted here (a year’s fit uses only fights before 1 January of that year and the seed Y, so cached rows stay valid); None fits every year. live_markets Group 6: the Kalshi / Polymarket columns at each round break (kalshi_, poly_), at most one row per (fight_id, round_idx), left-merged after the model features exactly as the research merged its two market frames.

Classes

RoundDataset

build

build(model_features: pd.DataFrame | None = None, live_markets: pd.DataFrame | None = None) -> dict[str, pd.DataFrame]

catalogue

catalogue(d2: pd.DataFrame, g1: list[str], g2: list[str], g3: list[str], g5: list[str], g6: list[str]) -> pd.DataFrame

eligible_fights

eligible_fights() -> pd.DataFrame

fight_frame

fight_frame(fights: pd.DataFrame, prof: pd.DataFrame, ctx: pd.DataFrame, pending: frozenset[str] = frozenset()) -> pd.DataFrame

One row per fight of fights (raw.bouts rows): keys, context, both fighters’ profiles, the matchup, the style records and the odds — everything dataset 1 holds per fight before the round rows. prof / ctx: FighterHistory.build() / CardContext.build(); pending: upcoming bouts (StyleSimilarity.pending).

fighter_panel

fighter_panel(prof: pd.DataFrame, fights: pd.DataFrame) -> pd.DataFrame

One row per fighter per eligible fight: his own profile and his own outcomes.

model_features

model_features(prof: pd.DataFrame, d1: pd.DataFrame, cached: pd.DataFrame | None = None) -> pd.DataFrame

Group 5, fitted walk-forward (the Bayesian curve takes ~15-60 s a year). cached rows are kept as they are; only the rows it lacks are fitted.

Functions

a_slot_of

a_slot_of(fight_id: pd.Series) -> pd.Series

1 or 2: which UFCStats slot is fighter A. Fixed per fight, independent of the result.