betting_combat.rounds.models.layer2
Layer 2: the round-start model on the live-market era (Kalshi), scored out of sample by purged combinatorial CV over cards, with the same fighter purge as layer 1.
Rows: Set 3 main (one per round start of a Kalshi-listed fight). Candidate feature groups
(symmetric only):
l1 the out-of-sample layer-1 predictions (CPCV, l1_*)
trade Set 3 main: the winner market’s price, book, volume and flow features
set1 the Dataset 1 compact feature set (context, matchup, odds, Bayesian clock)
stats the Dataset 2 layer-1 predictions (l1s_*) and the Group 3 round stats
side Set 3 side: the side markets (method of victory, distance, …), side_*
Targets, each starting from its best available prior (offset) and learning the correction:
y_win_fight prior: the live Kalshi price (mid_a)
y_decision prior: layer 1’s P(decision from here)
y_finish_round prior: layer 1’s P(finish this round)
The model: LightGBM on the residual (small, slow trees), bagged over 5 seeds.
nested the honest estimate: feature selection itself out of sample layer2_logistic P(ends this round) by the logistic family on the top-15 |IC| candidates, nested and purged (the final model’s second half)
Research: ufc/modeling/layer2.py (load, Layer2, nested) and
ufc/modeling/execution.py (layer2_logistic). Log-odds clip 0.5 % (logit_layer2).
Classes
Layer2
card_ci
card_ci(target: str, p_new: pd.Series, p_ref: pd.Series, n_boot: int = 2000, seed: int = 0) -> tuple[float, float, float, int]Log loss per row of p_new minus p_ref, with a 90 % interval from resampling
cards: (estimate, low, high, cards).
forward
forward(target: str, cand: list[str], max_features: int = FORWARD_MAX_FEATURES) -> tuple[list[str], pd.DataFrame]Forward selection: a candidate joins when the CV log loss drops by >= MIN_GAIN.
ic
ic(target: str, feats: list[str]) -> pd.DataFrameSpearman IC per feature: the mean over the CV’s time groups (those where the feature varies and the correlation is finite; at least 3), and its t.
mda
mda(target: str, feats: list[str], seed: int = 0) -> pd.SeriesPermutation importance out of sample: log-loss increase when a feature is shuffled within the test rows of every split (one fit per split, 1 seed for speed).
oos
oos(target: str, feats: list[str], rows: pd.Index | None = None, seeds: Sequence[int] = SEEDS) -> pd.SeriesOut-of-sample probability per row: bagged residual LightGBM on the target’s prior, averaged over the CPCV splits in which the row was a test row.
redundancy
redundancy(ic: pd.DataFrame, threshold: float = 0.9) -> list[str]score
score(target: str, p: pd.Series) -> floatFunctions
bagged_residual
bagged_residual(X_tr: pd.DataFrame, y_tr: pd.Series, z_tr: np.ndarray, X_te: pd.DataFrame) -> np.ndarrayRaw correction on the prior’s log-odds: LightGBM over SEEDS, averaged.
fit_bagged
fit_bagged(X_tr: pd.DataFrame, y_tr: pd.Series, z_tr: np.ndarray) -> tuple[lgb.LGBMClassifier, ...]bagged_residual’s models, one per seed of SEEDS (locked for later rows).
layer2_data
layer2_data(dataset1: pd.DataFrame, dataset2: pd.DataFrame, set3_main: pd.DataFrame, set3_side: pd.DataFrame, oos_layer1: pd.DataFrame, catalogue: pd.DataFrame, final_features: Sequence[str]) -> tuple[pd.DataFrame, dict[str, list[str]]]Layer 2’s rows (Set 3 main, joined to Dataset 1, the layer-1 CPCV predictions, the Group 3 round stats and Set 3 side) and its candidate feature groups.
catalogue: the dataset catalogue (column, group); final_features: the
Dataset 1 final compact feature list (eval/dataset1_final_features).
layer2_logistic
layer2_logistic(data: pd.DataFrame, groups: dict[str, list[str]]) -> pd.SeriesP(ends this round), logistic family, nested and purged, on the rows with side markets: per outer fold (6 groups, 1 out) the inner rows’ LOGISTIC_TOP largest-|IC| candidates of the l1 / trade / set1 / side groups, the prior’s log-odds as an input. Indexed by (fight_id, round_idx), in the outer folds’ order.
ll
ll(y: Any, p: Any) -> np.ndarraynested
nested(L: Layer2, target: str, groups: list[str], rows: pd.Index | None = None, top: int = NESTED_TOP, max_features: int = NESTED_MAX_FEATURES) -> pd.SeriesOut-of-sample probabilities where the feature selection itself is out of sample. Outer: leave one time group of cards out (6 folds, fighter purge). Inner, on the outer training cards only: rank candidates by IC, forward-select on a purged 5-group CV, then fit the bagged residual model and predict the held-out cards.
predict_bagged
predict_bagged(models: Sequence[lgb.LGBMClassifier], X_te: pd.DataFrame) -> np.ndarrayThe bagged models’ raw correction on new rows, averaged in seed order.
rank_by_ic
rank_by_ic(ic: pd.DataFrame) -> list[str]Features (ic’s index) by |IC|, largest first; equal |IC| in feature-name order (a
unique key: the ranking never depends on how a sort breaks ties).
ranked_by_ic
ranked_by_ic(inner: Layer2, target: str, cand: list[str]) -> list[str]Candidates by |IC| on inner’s rows, largest first (rank_by_ic).
select_forward
select_forward(inner: Layer2, target: str, cand: list[str], max_features: int = NESTED_MAX_FEATURES) -> list[str]The nested procedure’s inner step: forward selection on inner’s CV (1 seed).