Skip to content

betting_combat.rounds.models.joint

Layer 1: the full outcome of a fight (who x how x when) at every round start, built from priors and corrected by a small tree model.

Two parts:

round model at the start of round k (fight still going), five chances for that round: A by KO/TKO, A by submission, B by KO/TKO, B by submission, no finish. Starts from the prior P(finish in round k) = the Bayesian clock (m_S_*, m_p0) P(A is the finisher) = the pre-fight market (p_mkt_a) P(KO | finish) = the bookmakers’ KO share where priced, else the training rows’ league share and LightGBM learns only the correction (multiclass, init_score = log prior). decision model P(A wins | the fight goes to the judges). Starts from the market’s logit and learns the correction on training fights that went the distance.

Chaining the round model over rounds k, k+1, …, R gives every cell of the outcome grid (winner x method x round, plus winner on the cards), and from the grid every target: who wins, ends this round, goes to the judges, method. All consistent by construction.

JointRoundModel Dataset 1 (what is known before the fight) JointStatsModel Dataset 2: the current round sees the round stats so far; later rounds use the fitted Dataset 1 model (base) JointSplitModel Dataset 2 split by source: when and how from Dataset 1, who finishes and who wins on the cards from the round-stats model (the production layer 1s)

The feature lists are the research’s final compact sets (eval/dataset{1,2}_final_features) minus the market price, which enters as the prior (compact_features).

Research: ufc/modeling/joint.py (models, explained, paired_ci) and ufc/modeling/cpcv.py (_cells). Log-odds clip 1 % (logit_joint).

Classes

JointRoundModel

X

X(d: pd.DataFrame) -> pd.DataFrame

choose_trees

choose_trees(train: pd.DataFrame) -> pd.DataFrame

Walk-forward inside train (fit < Y, score Y, Y = 2016-2019), per tree count and part: mean out-of-sample log loss and its standard error across the folds. The one-standard-error rule picks the fewest trees within one SE of the best (sets n_round and n_dec; ko_share_league is left at the last fold’s).

curve

curve(d: pd.DataFrame) -> np.ndarray

(rows x 6) S at the end of rounds 0..5: S(0) = 1, S(final bell) = 0 and after.

fit

fit(train: pd.DataFrame) -> JointRoundModel

grid

grid(d: pd.DataFrame) -> dict[str, np.ndarray]

For every row (alive at the start of round r): P of every cell from here on. Returns arrays (rows x 5 rounds) for a_ko, a_sub, b_ko, b_sub, and vectors dec, dec_a.

hazard

hazard(d: pd.DataFrame, k: np.ndarray) -> np.ndarray

P(finish in round k+1 | alive at its start) = (1-p0)(S_k - S_k+1) / (p0 + (1-p0) S_k).

prior

prior(d: pd.DataFrame) -> np.ndarray

round_label

round_label(d: pd.DataFrame) -> np.ndarray

targets

targets(d: pd.DataFrame) -> pd.DataFrame

JointSplitModel

Bases: JointStatsModel

Dataset 2, split by what each source is good at (validation 2020-2022): whether the round ends and how -> the Dataset 1 round model who is the finisher -> the round-stats model who wins on the cards -> the round-stats decision model

grid

grid(d: pd.DataFrame) -> dict[str, np.ndarray]

JointStatsModel

Bases: JointRoundModel

The same grid with round stats (Dataset 2), where they can be known:

current round the round model sees the stats of the rounds already fought later rounds the Dataset 1 round model (their stats cannot be known yet) on the cards P(A wins | decision) sees the stats so far (who is ahead), trained on every round start of fights that went the distance

grid

grid(d: pd.DataFrame) -> dict[str, np.ndarray]

Functions

cells

cells(model: JointRoundModel, d: pd.DataFrame) -> pd.DataFrame

The layer-1 columns per row (from here to the end of the fight): the five targets and the six outcome cells a/b x ko/sub/dec.

compact_features

compact_features(final_features: Iterable[str]) -> list[str]

The joint model’s features: the research’s final compact set without the market price (the market enters as the prior).

explained

explained(y: Any, p: Any, labels: Sequence[int] = (0, 1)) -> float

Share of log loss explained against the base rate (binary: the mean; multiclass: the class frequencies).

paired_ci

paired_ci(d: pd.DataFrame, model: JointRoundModel, n_boot: int = 2000, seed: int = 0) -> pd.DataFrame

Joint model minus benchmark, log loss per row, with a 90% interval from resampling whole cards (events): fights on one card are not independent. Negative = joint better.

softmax

softmax(z: np.ndarray) -> np.ndarray