betting_combat.rounds.features.dataset
RoundDataset: assemble the three tables and their column catalogue.
Port of the research ufc/dataset/build.py; the numbers must stay identical. The
research wrote files; production returns the tables and the caller stores them.
dataset1 one row per (fight, round start); keys + context + Group 1 (both fighters) + Group 2 (matchup) + Group 4 (pre-fight odds) + targets + Group 5 (clock model) + Group 6 (live markets, when given)dataset2 exactly dataset1's rows and columns + Group 3 (rounds fought so far)fighters one row per (fighter, fight): that fighter's Group 1 profile alone and fighter-level targets (win this fight, win this and the next, finish, get finished)catalogue every column: its group, and when it is known (pre-fight / live at the round break from Kalshi / live only with a stats feed / outcome)model_features Group 5 as used (the research's cache, extended where rows were missing)Rows kept: fights from 2008 on, won or lost (draws, no contests, DQs, overturned and ‘other’ results removed), 3- or 5-round fights of 5-minute rounds, with round stats for every round that started. Fighter histories use every fight back to 1994.
A/B: fighter A is UFCStats slot 1 or 2 by a fixed hash of the fight id (UFCStats lists the winner first 63% of the time, so the slots must not be used as they come).
Two inputs come from elsewhere:
model_features the Group 5 cache (clock.ClockFeatures output). Rows it lacks are
fitted here (a year’s fit uses only fights before 1 January of that year
and the seed Y, so cached rows stay valid); None fits every year.
live_markets Group 6: the Kalshi / Polymarket columns at each round break (kalshi_,
poly_), at most one row per (fight_id, round_idx), left-merged after the
model features exactly as the research merged its two market frames.
Classes
RoundDataset
build
build(model_features: pd.DataFrame | None = None, live_markets: pd.DataFrame | None = None) -> dict[str, pd.DataFrame]catalogue
catalogue(d2: pd.DataFrame, g1: list[str], g2: list[str], g3: list[str], g5: list[str], g6: list[str]) -> pd.DataFrameeligible_fights
eligible_fights() -> pd.DataFramefight_frame
fight_frame(fights: pd.DataFrame, prof: pd.DataFrame, ctx: pd.DataFrame, pending: frozenset[str] = frozenset()) -> pd.DataFrameOne row per fight of fights (raw.bouts rows): keys, context, both fighters’
profiles, the matchup, the style records and the odds — everything dataset 1 holds per
fight before the round rows. prof / ctx: FighterHistory.build() /
CardContext.build(); pending: upcoming bouts (StyleSimilarity.pending).
fighter_panel
fighter_panel(prof: pd.DataFrame, fights: pd.DataFrame) -> pd.DataFrameOne row per fighter per eligible fight: his own profile and his own outcomes.
model_features
model_features(prof: pd.DataFrame, d1: pd.DataFrame, cached: pd.DataFrame | None = None) -> pd.DataFrameGroup 5, fitted walk-forward (the Bayesian curve takes ~15-60 s a year). cached
rows are kept as they are; only the rows it lacks are fitted.
Functions
a_slot_of
a_slot_of(fight_id: pd.Series) -> pd.Series1 or 2: which UFCStats slot is fighter A. Fixed per fight, independent of the result.