betting_combat.rounds.training
The rounds strategy’s training pipeline (the weekly refit) and the forward lock-and-score.
REFIT (every step out of sample; the outputs feed the next step)
layer1_oos layer 1 by purged CPCV over cards (10 time groups, 2 held out: 45
splits, every row tested 9 times, predictions averaged): the Dataset 1
joint model (50 trees, decision part = the market; l1_*) and the
Dataset 2 split model on top of it (150 / 150 trees; l1s_*)
model_probs layer 2, nested: P(ends this round) on the rows with side markets
(finish_side), P(A wins) from the core (win_core), and layer 1’s
columns the trading step uses
final_events the final P(ends this round) = mean of the bagged boosted trees
(finish_side) and the logistic family (layer2.layer2_logistic),
layer 1 where a fight has no side markets; the candidate bets; their
uniqueness weights; the meta probability = mean of logistic, extra trees
and bagged boosted trees, each by purged CPCV
FORWARD (forward): lock every model on the cards before a cutoff date, then score the
cards on or after it exactly as production would. Nothing from the scored cards is used
to fit anything: the training inputs (layer 1’s CPCV predictions, layer 2’s nested
probabilities, the candidate bets and their meta features) are re-derived from the rows
before the cutoff only (forward_inputs; research fixed 2026-09-30), and forward
refuses inputs that are not (check_before).
layer 1 JointRoundModel (50 trees, decision part = the market), fitted on every round
start before the cutoff that has a pre-fight price
layer 2 P(ends this round) on rows with side markets: the mean of bagged boosted trees
on features forward-selected on the training rows (purged CV over cards inside
them) and the logistic model on the training rows’ 15 largest-|IC| candidates;
trained on the out-of-sample layer-1 predictions, applied with the locked
layer-1 model’s predictions for the scored cards
distance (1 - layer 2) x layer 1’s no-finish chance in every later round
meta logistic + extra trees + bagged boosted trees, fitted on every candidate bet
before the cutoff (uniqueness weights), averaged; the same features as
final_events (layer 2’s own probability l2_p_finish included)
Data arrive as arguments (no IO here): Dataset 1 / 2 (round starts), layer 2’s rows and
groups (layer2.layer2_data), the Kalshi fight clock (fights: fight_id,
start, snap, …) and the Set 4 minute books (minutes).
Research: ufc/modeling/cpcv.py::layer1_oos, ufc/modeling/flat_bets.py::model_probs,
ufc/modeling/execution.py::final_events, ufc/modeling/forward.py (without the
backtest).
Classes
ForwardInputs
A forward test’s training inputs, derived from the rows before its cutoff only.
ForwardScore
The locked models’ outputs on the scored cards.
LockedLayer2
Layer 2 locked on the side-market rows before a cutoff: the bagged boosted trees on
the forward-selected features (kept; none selected -> layer 1 alone) and the
logistic family on the top-|IC| candidates (top), both on layer 1’s P(ends this
round) as the prior.
score
score(rows: pd.DataFrame) -> pd.DataFramefinish_trees, finish_logistic and their mean finish_final for side-market
rows (layer 2’s columns, l1_p_finish_round from the locked layer 1).
LockedMeta
The meta model locked on the candidate bets before a cutoff: the META3 families
fitted with uniqueness weights on features; P(bet wins) = their mean.
score
score(bets: pd.DataFrame) -> np.ndarrayLockedModels
Every model the rounds strategy trades with, locked on the cards before cutoff
(the weekly refit’s product; forward scores later cards with it).
Functions
check_before
check_before(cutoff: pd.Timestamp, **frames: pd.DataFrame) -> NoneRefuse training inputs with a row dated on or after cutoff (event_date).
check_layer1_before
check_layer1_before(cutoff: pd.Timestamp, data: pd.DataFrame, groups: dict[str, list[str]]) -> NoneRefuse layer-2 rows whose layer-1 CPCV predictions were derived with the rows on or
after cutoff in the CV (those rows then carry predictions too).
clocked
clocked(d: pd.DataFrame) -> pd.DataFrameThe round starts the Bayesian clock prices (m_S_1 and m_p0 known).
final_events
final_events(probs: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], fights: pd.DataFrame, minutes: pd.DataFrame) -> pd.DataFrameEvery candidate bet on the final model with its meta features, uniqueness weight
(weight), the out-of-sample meta probability of each META3 family
(meta_logistic, …) and their mean (meta).
final_probs
final_probs(probs: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]]) -> pd.DataFrameprobs (model_probs) with the logistic layer 2 (finish_logistic) and the
final P(ends this round): the mean of the two layer 2s where present, else layer 1
(finish_final).
fit_layer1
fit_layer1(dataset1: pd.DataFrame, cutoff: pd.Timestamp, features1: Sequence[str]) -> JointRoundModelLayer 1 locked on the clocked round starts before cutoff with a pre-fight price.
fit_layer2
fit_layer2(cutoff: pd.Timestamp, data: pd.DataFrame, groups: dict[str, list[str]]) -> LockedLayer2Layer 2 locked on the side-market rows of data before cutoff (their layer-1
columns are the out-of-sample CPCV predictions layer 2 is trained on).
fit_locked
fit_locked(cutoff: pd.Timestamp, dataset1: pd.DataFrame, features1: Sequence[str], data: pd.DataFrame, groups: dict[str, list[str]], events: pd.DataFrame, fights: pd.DataFrame, drop_l2_p_finish: bool = False, check: bool = True) -> LockedModelsLock every model on the cards before cutoff (the weekly refit’s product): layer 1
on Dataset 1, layer 2 on the refit’s layer-2 rows (data / groups, with the CPCV
layer-1 columns), the meta model on the refit’s candidate bets (events,
final_events). Only rows before cutoff are fitted on, and every derived training
input must itself predate it (check_before, check_layer1_before: a CPCV or a
nested prediction that saw later cards is refused, as forward refuses it).
check=False only to reproduce a reference fitted from the full-period refit.
fit_meta
fit_meta(train_events: pd.DataFrame, groups: dict[str, list[str]], fights: pd.DataFrame, drop_l2_p_finish: bool = False) -> LockedMetaMETA3 fitted on the training events (final_events rows, which carry their meta
features), weighted by their average uniqueness.
drop_l2_p_finish: fit without layer 2’s own probability. The research forward test
did until 2026-09-30 (a research bug: under pandas 2 its second features merge split
the column into l2_p_finish_x / _y; fixed there since); production and the
current references keep it (False).
forward
forward(cutoff: pd.Timestamp, dataset1: pd.DataFrame, features1: Sequence[str], data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame, events: pd.DataFrame, fights: pd.DataFrame, minutes: pd.DataFrame, drop_l2_p_finish: bool = False) -> ForwardScoreLock every model on the cards before cutoff (fit_locked) and score the cards
on or after it: the candidate bets with their meta probability (events.meta).
data / groups / probs / events: layer 2’s rows, model_probs and the
candidate bets with their meta features, all derived from the rows before cutoff
only (forward_inputs; refused otherwise: check_before, check_layer1_before).
drop_l2_p_finish: see fit_meta (production and the research: False).
forward_events
forward_events(cutoff: pd.Timestamp, dataset1: pd.DataFrame, l1_test: pd.DataFrame, l2: pd.DataFrame, fights: pd.DataFrame, minutes: pd.DataFrame) -> pd.DataFrameThe candidate bets on the scored cards from the locked models (layer 2 where the row has side markets, else layer 1).
forward_inputs
forward_inputs(cutoff: pd.Timestamp, dataset1: pd.DataFrame, dataset2: pd.DataFrame, features1: Sequence[str], features2: Sequence[str], set3_main: pd.DataFrame, set3_side: pd.DataFrame, catalogue: pd.DataFrame, final_features: Sequence[str], fights: pd.DataFrame, minutes: pd.DataFrame) -> ForwardInputsThe refit (layer1_oos, model_probs, training_events) on the rows before
cutoff only, as the research forward test derives its training inputs
(forward.pre_cutoff). features1 / features2: the joint models’ features;
final_features: the Dataset 1 final compact list (layer2_data).
forward_layer1
forward_layer1(dataset1: pd.DataFrame, cutoff: pd.Timestamp, features1: Sequence[str]) -> tuple[JointRoundModel, pd.DataFrame]Layer 1 locked on the round starts before cutoff (with a pre-fight price) and its
columns (l1_*) on every clocked round start on or after it.
forward_layer2
forward_layer2(cutoff: pd.Timestamp, l1_test: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], locked: LockedLayer2 | None = None) -> pd.DataFrameLayer 2 locked on the side-market rows before cutoff: the bagged trees on the
features forward-selected there, the logistic model on the top-|IC| candidates, and
their mean, on the side-market rows on or after it (whose layer-1 columns are the
locked layer 1’s). attrs: kept and top. locked: the fitted layer 2
(fit_layer2), fitted here when not given.
forward_meta
forward_meta(train_events: pd.DataFrame, test_events: pd.DataFrame, l1_test: pd.DataFrame, l2: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame, fights: pd.DataFrame, drop_l2_p_finish: bool = False, locked: LockedMeta | None = None) -> pd.SeriesThe meta probability of every scored candidate: META3 fitted on the training events
(uniqueness weights), averaged (fit_meta, unless locked is given). The scored
rows’ layer-1 columns and layer-2 input are the locked models’ (meta_rows).
train_events: training_events / final_events rows (they carry their meta
features; used as they are, never merged a second time). Until 2026-09-30 the research
ran features on them again, which under pandas 2 split l2_p_finish into
l2_p_finish_x / _y and so fitted its forward meta models WITHOUT it
(drop_l2_p_finish=True reproduces that); the research now fits with it, as
production does.
layer1_oos
layer1_oos(dataset1: pd.DataFrame, dataset2: pd.DataFrame, features1: Sequence[str], features2: Sequence[str]) -> tuple[pd.DataFrame, pd.DataFrame]Layer 1’s out-of-sample predictions by purged CPCV, one row per (fight_id, round_idx)
the clock prices: the mean over the splits in which the row was a test row of the
Dataset 1 joint model’s columns (l1_) and the Dataset 2 split model’s (l1s_);
l1_n_splits = how many splits scored the row, cv_group its time group.
Also returns the split log (test groups, rows, purged fights).
features1 / features2: the joint models’ features (joint.compact_features
of the Dataset 1 / 2 final lists).
meta_rows
meta_rows(test_events: pd.DataFrame, l1_test: pd.DataFrame, l2: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame) -> pd.DataFrameThe scored candidates joined to their meta features: the scored cards’ layer-1
columns and layer-2 input are the locked models’ (the bagged trees’ finish_trees,
as in training).
model_probs
model_probs(data: pd.DataFrame, groups: dict[str, list[str]]) -> pd.DataFrameOut-of-sample probabilities per round-start row (layer 2’s rows): finish_side P(ends this round): nested layer 2 on the core + side markets (rows with side markets; NaN elsewhere) finish_l1 P(ends this round): layer 1 (CPCV), every row decision_l1 P(goes the distance from here): layer 1 (CPCV) win_core P(A wins): nested layer 2 on the core, starting from the live price a_finisher P(A is the finisher | a finish): layer 1 (CPCV) grid cells and the targets.
training_events
training_events(probs: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], fights: pd.DataFrame, minutes: pd.DataFrame) -> tuple[pd.DataFrame, list[str], Any]Every candidate bet on the final model with its meta features and uniqueness weight
(weight): final_events without the meta probabilities. Returns (events, the meta
feature list, the events’ indicator matrix).