Skip to content

betting_combat.rounds.training

The rounds strategy’s training pipeline (the weekly refit) and the forward lock-and-score.

REFIT (every step out of sample; the outputs feed the next step) layer1_oos layer 1 by purged CPCV over cards (10 time groups, 2 held out: 45 splits, every row tested 9 times, predictions averaged): the Dataset 1 joint model (50 trees, decision part = the market; l1_*) and the Dataset 2 split model on top of it (150 / 150 trees; l1s_*) model_probs layer 2, nested: P(ends this round) on the rows with side markets (finish_side), P(A wins) from the core (win_core), and layer 1’s columns the trading step uses final_events the final P(ends this round) = mean of the bagged boosted trees (finish_side) and the logistic family (layer2.layer2_logistic), layer 1 where a fight has no side markets; the candidate bets; their uniqueness weights; the meta probability = mean of logistic, extra trees and bagged boosted trees, each by purged CPCV

FORWARD (forward): lock every model on the cards before a cutoff date, then score the cards on or after it exactly as production would. Nothing from the scored cards is used to fit anything: the training inputs (layer 1’s CPCV predictions, layer 2’s nested probabilities, the candidate bets and their meta features) are re-derived from the rows before the cutoff only (forward_inputs; research fixed 2026-09-30), and forward refuses inputs that are not (check_before). layer 1 JointRoundModel (50 trees, decision part = the market), fitted on every round start before the cutoff that has a pre-fight price layer 2 P(ends this round) on rows with side markets: the mean of bagged boosted trees on features forward-selected on the training rows (purged CV over cards inside them) and the logistic model on the training rows’ 15 largest-|IC| candidates; trained on the out-of-sample layer-1 predictions, applied with the locked layer-1 model’s predictions for the scored cards distance (1 - layer 2) x layer 1’s no-finish chance in every later round meta logistic + extra trees + bagged boosted trees, fitted on every candidate bet before the cutoff (uniqueness weights), averaged; the same features as final_events (layer 2’s own probability l2_p_finish included)

Data arrive as arguments (no IO here): Dataset 1 / 2 (round starts), layer 2’s rows and groups (layer2.layer2_data), the Kalshi fight clock (fights: fight_id, start, snap, …) and the Set 4 minute books (minutes).

Research: ufc/modeling/cpcv.py::layer1_oos, ufc/modeling/flat_bets.py::model_probs, ufc/modeling/execution.py::final_events, ufc/modeling/forward.py (without the backtest).

Classes

ForwardInputs

A forward test’s training inputs, derived from the rows before its cutoff only.

ForwardScore

The locked models’ outputs on the scored cards.

LockedLayer2

Layer 2 locked on the side-market rows before a cutoff: the bagged boosted trees on the forward-selected features (kept; none selected -> layer 1 alone) and the logistic family on the top-|IC| candidates (top), both on layer 1’s P(ends this round) as the prior.

score

score(rows: pd.DataFrame) -> pd.DataFrame

finish_trees, finish_logistic and their mean finish_final for side-market rows (layer 2’s columns, l1_p_finish_round from the locked layer 1).

LockedMeta

The meta model locked on the candidate bets before a cutoff: the META3 families fitted with uniqueness weights on features; P(bet wins) = their mean.

score

score(bets: pd.DataFrame) -> np.ndarray

LockedModels

Every model the rounds strategy trades with, locked on the cards before cutoff (the weekly refit’s product; forward scores later cards with it).

Functions

check_before

check_before(cutoff: pd.Timestamp, **frames: pd.DataFrame) -> None

Refuse training inputs with a row dated on or after cutoff (event_date).

check_layer1_before

check_layer1_before(cutoff: pd.Timestamp, data: pd.DataFrame, groups: dict[str, list[str]]) -> None

Refuse layer-2 rows whose layer-1 CPCV predictions were derived with the rows on or after cutoff in the CV (those rows then carry predictions too).

clocked

clocked(d: pd.DataFrame) -> pd.DataFrame

The round starts the Bayesian clock prices (m_S_1 and m_p0 known).

final_events

final_events(probs: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], fights: pd.DataFrame, minutes: pd.DataFrame) -> pd.DataFrame

Every candidate bet on the final model with its meta features, uniqueness weight (weight), the out-of-sample meta probability of each META3 family (meta_logistic, …) and their mean (meta).

final_probs

final_probs(probs: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]]) -> pd.DataFrame

probs (model_probs) with the logistic layer 2 (finish_logistic) and the final P(ends this round): the mean of the two layer 2s where present, else layer 1 (finish_final).

fit_layer1

fit_layer1(dataset1: pd.DataFrame, cutoff: pd.Timestamp, features1: Sequence[str]) -> JointRoundModel

Layer 1 locked on the clocked round starts before cutoff with a pre-fight price.

fit_layer2

fit_layer2(cutoff: pd.Timestamp, data: pd.DataFrame, groups: dict[str, list[str]]) -> LockedLayer2

Layer 2 locked on the side-market rows of data before cutoff (their layer-1 columns are the out-of-sample CPCV predictions layer 2 is trained on).

fit_locked

fit_locked(cutoff: pd.Timestamp, dataset1: pd.DataFrame, features1: Sequence[str], data: pd.DataFrame, groups: dict[str, list[str]], events: pd.DataFrame, fights: pd.DataFrame, drop_l2_p_finish: bool = False, check: bool = True) -> LockedModels

Lock every model on the cards before cutoff (the weekly refit’s product): layer 1 on Dataset 1, layer 2 on the refit’s layer-2 rows (data / groups, with the CPCV layer-1 columns), the meta model on the refit’s candidate bets (events, final_events). Only rows before cutoff are fitted on, and every derived training input must itself predate it (check_before, check_layer1_before: a CPCV or a nested prediction that saw later cards is refused, as forward refuses it). check=False only to reproduce a reference fitted from the full-period refit.

fit_meta

fit_meta(train_events: pd.DataFrame, groups: dict[str, list[str]], fights: pd.DataFrame, drop_l2_p_finish: bool = False) -> LockedMeta

META3 fitted on the training events (final_events rows, which carry their meta features), weighted by their average uniqueness.

drop_l2_p_finish: fit without layer 2’s own probability. The research forward test did until 2026-09-30 (a research bug: under pandas 2 its second features merge split the column into l2_p_finish_x / _y; fixed there since); production and the current references keep it (False).

forward

forward(cutoff: pd.Timestamp, dataset1: pd.DataFrame, features1: Sequence[str], data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame, events: pd.DataFrame, fights: pd.DataFrame, minutes: pd.DataFrame, drop_l2_p_finish: bool = False) -> ForwardScore

Lock every model on the cards before cutoff (fit_locked) and score the cards on or after it: the candidate bets with their meta probability (events.meta).

data / groups / probs / events: layer 2’s rows, model_probs and the candidate bets with their meta features, all derived from the rows before cutoff only (forward_inputs; refused otherwise: check_before, check_layer1_before). drop_l2_p_finish: see fit_meta (production and the research: False).

forward_events

forward_events(cutoff: pd.Timestamp, dataset1: pd.DataFrame, l1_test: pd.DataFrame, l2: pd.DataFrame, fights: pd.DataFrame, minutes: pd.DataFrame) -> pd.DataFrame

The candidate bets on the scored cards from the locked models (layer 2 where the row has side markets, else layer 1).

forward_inputs

forward_inputs(cutoff: pd.Timestamp, dataset1: pd.DataFrame, dataset2: pd.DataFrame, features1: Sequence[str], features2: Sequence[str], set3_main: pd.DataFrame, set3_side: pd.DataFrame, catalogue: pd.DataFrame, final_features: Sequence[str], fights: pd.DataFrame, minutes: pd.DataFrame) -> ForwardInputs

The refit (layer1_oos, model_probs, training_events) on the rows before cutoff only, as the research forward test derives its training inputs (forward.pre_cutoff). features1 / features2: the joint models’ features; final_features: the Dataset 1 final compact list (layer2_data).

forward_layer1

forward_layer1(dataset1: pd.DataFrame, cutoff: pd.Timestamp, features1: Sequence[str]) -> tuple[JointRoundModel, pd.DataFrame]

Layer 1 locked on the round starts before cutoff (with a pre-fight price) and its columns (l1_*) on every clocked round start on or after it.

forward_layer2

forward_layer2(cutoff: pd.Timestamp, l1_test: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], locked: LockedLayer2 | None = None) -> pd.DataFrame

Layer 2 locked on the side-market rows before cutoff: the bagged trees on the features forward-selected there, the logistic model on the top-|IC| candidates, and their mean, on the side-market rows on or after it (whose layer-1 columns are the locked layer 1’s). attrs: kept and top. locked: the fitted layer 2 (fit_layer2), fitted here when not given.

forward_meta

forward_meta(train_events: pd.DataFrame, test_events: pd.DataFrame, l1_test: pd.DataFrame, l2: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame, fights: pd.DataFrame, drop_l2_p_finish: bool = False, locked: LockedMeta | None = None) -> pd.Series

The meta probability of every scored candidate: META3 fitted on the training events (uniqueness weights), averaged (fit_meta, unless locked is given). The scored rows’ layer-1 columns and layer-2 input are the locked models’ (meta_rows).

train_events: training_events / final_events rows (they carry their meta features; used as they are, never merged a second time). Until 2026-09-30 the research ran features on them again, which under pandas 2 split l2_p_finish into l2_p_finish_x / _y and so fitted its forward meta models WITHOUT it (drop_l2_p_finish=True reproduces that); the research now fits with it, as production does.

layer1_oos

layer1_oos(dataset1: pd.DataFrame, dataset2: pd.DataFrame, features1: Sequence[str], features2: Sequence[str]) -> tuple[pd.DataFrame, pd.DataFrame]

Layer 1’s out-of-sample predictions by purged CPCV, one row per (fight_id, round_idx) the clock prices: the mean over the splits in which the row was a test row of the Dataset 1 joint model’s columns (l1_) and the Dataset 2 split model’s (l1s_); l1_n_splits = how many splits scored the row, cv_group its time group. Also returns the split log (test groups, rows, purged fights).

features1 / features2: the joint models’ features (joint.compact_features of the Dataset 1 / 2 final lists).

meta_rows

meta_rows(test_events: pd.DataFrame, l1_test: pd.DataFrame, l2: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame) -> pd.DataFrame

The scored candidates joined to their meta features: the scored cards’ layer-1 columns and layer-2 input are the locked models’ (the bagged trees’ finish_trees, as in training).

model_probs

model_probs(data: pd.DataFrame, groups: dict[str, list[str]]) -> pd.DataFrame

Out-of-sample probabilities per round-start row (layer 2’s rows): finish_side P(ends this round): nested layer 2 on the core + side markets (rows with side markets; NaN elsewhere) finish_l1 P(ends this round): layer 1 (CPCV), every row decision_l1 P(goes the distance from here): layer 1 (CPCV) win_core P(A wins): nested layer 2 on the core, starting from the live price a_finisher P(A is the finisher | a finish): layer 1 (CPCV) grid cells and the targets.

training_events

training_events(probs: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], fights: pd.DataFrame, minutes: pd.DataFrame) -> tuple[pd.DataFrame, list[str], Any]

Every candidate bet on the final model with its meta features and uniqueness weight (weight): final_events without the meta probabilities. Returns (events, the meta feature list, the events’ indicator matrix).