betting_combat.rounds.models.meta
Meta-labelling (Lopez de Prado, AFML ch. 3): the primary model (final layer 1 + 2) picks
the side of every candidate bet; the meta model learns P(this bet wins), so whether to take
it and how big (common.afml_size).
FEATURES (all out of sample for every row): the core (Set 1 compact, Set 3 trade, Set 3
side), layer 1 probabilities (CPCV), layer 2’s P(ends this round) (nested CV), the
final probability for the market, and the bet itself: market, side, edge, entry
price, spread, round.
WEIGHTS Lopez average uniqueness over each bet’s life (entry -> the market prices the
result), counted per fight (ml.lopez).
MODELS logistic (L2, standardised, medians for gaps), random forest / extra trees
(balanced subsample class weights), sequential-bootstrap trees, bagged small
LightGBM; the production meta is the mean of logistic, extra trees and bagged
LightGBM (META3). Out of sample: purged combinatorial CV over cards (6 groups,
2 held out; fighter purge), each event’s probability the mean over the splits where
it was held out.
The fight clock (fights) is a table with one row per Kalshi-listed fight: fight_id,
start (the estimated start, UTC) and snap (the minute the market priced the
result), among others.
Research: ufc/modeling/meta.py (add_lifetimes, features, fit_predict,
oos_family, size_from_prob, FAMILIES, TREE).
Classes
FittedMeta
One fitted meta family (fit_predict split in two, so a model locked on the
training events can price new bets later): the model(s), the training medians that fill
gaps (not used by the bagged LightGBM, which takes gaps as they are) and the features.
Functions
add_lifetimes
add_lifetimes(x: pd.DataFrame, fights: pd.DataFrame) -> pd.DataFrameEach bet’s life: from entry (the row’s time) to the moment the market priced the
result. Adds t_entry and t_end to x in place and returns it.
core_features
core_features(groups: dict[str, list[str]]) -> list[str]The core the meta model sees: layer 1, Set 3 trade, Set 1 compact, Set 3 side.
entry_time
entry_time(start: pd.Timestamp, round_idx: int) -> pd.TimestampThe decision minute of a round start: round 1, the minute before the fight; round k + 1, the end of the break (start + 6 k minutes), floored to the minute.
feature_list
feature_list(x: pd.DataFrame, groups: dict[str, list[str]]) -> list[str]The meta features present in x: the bet’s own, then the core; numeric and varying.
features
features(ev: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame) -> tuple[pd.DataFrame, list[str]]The events joined to their meta features, and the feature list (numeric, varying).
data / groups: layer 2’s rows and feature groups (layer2.layer2_data);
probs: the out-of-sample model probabilities per round start (finish_side, the
nested layer-2 P(ends this round), becomes l2_p_finish).
fit
fit(kind: str, Xtr: pd.DataFrame, ytr: pd.Series, wtr: pd.Series, ind_tr: sparse.csr_matrix) -> FittedMetaFit meta family kind on the training events (see fit_predict).
fit_predict
fit_predict(kind: str, Xtr: pd.DataFrame, ytr: pd.Series, wtr: pd.Series, Xte: pd.DataFrame, ind_tr: sparse.csr_matrix) -> np.ndarrayFit meta family kind on the training events (weights wtr; ind_tr their
indicator matrix, for the sequential bootstrap) and return P(bet wins) on Xte.
oos_family
oos_family(x: pd.DataFrame, feats: list[str], kind: str, ind: sparse.csr_matrix) -> pd.SeriesOut-of-sample P(bet wins) per event (label won, weights weight): the mean over
the purged CPCV splits in which the event was held out. ind: the events’ indicator
matrix, rows in x’s order.
predict
predict(fitted: FittedMeta, Xte: pd.DataFrame) -> np.ndarrayP(bet wins) for new bets (Xte: the fitted features, in order).
size_from_prob
size_from_prob(m: pd.Series) -> pd.SeriesAFML 10.1: bet size from the meta probability; nothing at or below 0.5.