Skip to content

betting_combat.rounds.models.meta

Meta-labelling (Lopez de Prado, AFML ch. 3): the primary model (final layer 1 + 2) picks the side of every candidate bet; the meta model learns P(this bet wins), so whether to take it and how big (common.afml_size).

FEATURES (all out of sample for every row): the core (Set 1 compact, Set 3 trade, Set 3 side), layer 1 probabilities (CPCV), layer 2’s P(ends this round) (nested CV), the final probability for the market, and the bet itself: market, side, edge, entry price, spread, round. WEIGHTS Lopez average uniqueness over each bet’s life (entry -> the market prices the result), counted per fight (ml.lopez). MODELS logistic (L2, standardised, medians for gaps), random forest / extra trees (balanced subsample class weights), sequential-bootstrap trees, bagged small LightGBM; the production meta is the mean of logistic, extra trees and bagged LightGBM (META3). Out of sample: purged combinatorial CV over cards (6 groups, 2 held out; fighter purge), each event’s probability the mean over the splits where it was held out.

The fight clock (fights) is a table with one row per Kalshi-listed fight: fight_id, start (the estimated start, UTC) and snap (the minute the market priced the result), among others.

Research: ufc/modeling/meta.py (add_lifetimes, features, fit_predict, oos_family, size_from_prob, FAMILIES, TREE).

Classes

FittedMeta

One fitted meta family (fit_predict split in two, so a model locked on the training events can price new bets later): the model(s), the training medians that fill gaps (not used by the bagged LightGBM, which takes gaps as they are) and the features.

Functions

add_lifetimes

add_lifetimes(x: pd.DataFrame, fights: pd.DataFrame) -> pd.DataFrame

Each bet’s life: from entry (the row’s time) to the moment the market priced the result. Adds t_entry and t_end to x in place and returns it.

core_features

core_features(groups: dict[str, list[str]]) -> list[str]

The core the meta model sees: layer 1, Set 3 trade, Set 1 compact, Set 3 side.

entry_time

entry_time(start: pd.Timestamp, round_idx: int) -> pd.Timestamp

The decision minute of a round start: round 1, the minute before the fight; round k + 1, the end of the break (start + 6 k minutes), floored to the minute.

feature_list

feature_list(x: pd.DataFrame, groups: dict[str, list[str]]) -> list[str]

The meta features present in x: the bet’s own, then the core; numeric and varying.

features

features(ev: pd.DataFrame, data: pd.DataFrame, groups: dict[str, list[str]], probs: pd.DataFrame) -> tuple[pd.DataFrame, list[str]]

The events joined to their meta features, and the feature list (numeric, varying).

data / groups: layer 2’s rows and feature groups (layer2.layer2_data); probs: the out-of-sample model probabilities per round start (finish_side, the nested layer-2 P(ends this round), becomes l2_p_finish).

fit

fit(kind: str, Xtr: pd.DataFrame, ytr: pd.Series, wtr: pd.Series, ind_tr: sparse.csr_matrix) -> FittedMeta

Fit meta family kind on the training events (see fit_predict).

fit_predict

fit_predict(kind: str, Xtr: pd.DataFrame, ytr: pd.Series, wtr: pd.Series, Xte: pd.DataFrame, ind_tr: sparse.csr_matrix) -> np.ndarray

Fit meta family kind on the training events (weights wtr; ind_tr their indicator matrix, for the sequential bootstrap) and return P(bet wins) on Xte.

oos_family

oos_family(x: pd.DataFrame, feats: list[str], kind: str, ind: sparse.csr_matrix) -> pd.Series

Out-of-sample P(bet wins) per event (label won, weights weight): the mean over the purged CPCV splits in which the event was held out. ind: the events’ indicator matrix, rows in x’s order.

predict

predict(fitted: FittedMeta, Xte: pd.DataFrame) -> np.ndarray

P(bet wins) for new bets (Xte: the fitted features, in order).

size_from_prob

size_from_prob(m: pd.Series) -> pd.Series

AFML 10.1: bet size from the meta probability; nothing at or below 0.5.