betting_combat.rounds.models.joint
Layer 1: the full outcome of a fight (who x how x when) at every round start, built from priors and corrected by a small tree model.
Two parts:
round model at the start of round k (fight still going), five chances for that round: A by KO/TKO, A by submission, B by KO/TKO, B by submission, no finish. Starts from the prior P(finish in round k) = the Bayesian clock (m_S_*, m_p0) P(A is the finisher) = the pre-fight market (p_mkt_a) P(KO | finish) = the bookmakers’ KO share where priced, else the training rows’ league share and LightGBM learns only the correction (multiclass, init_score = log prior). decision model P(A wins | the fight goes to the judges). Starts from the market’s logit and learns the correction on training fights that went the distance.
Chaining the round model over rounds k, k+1, …, R gives every cell of the outcome grid (winner x method x round, plus winner on the cards), and from the grid every target: who wins, ends this round, goes to the judges, method. All consistent by construction.
JointRoundModel Dataset 1 (what is known before the fight)
JointStatsModel Dataset 2: the current round sees the round stats so far; later rounds
use the fitted Dataset 1 model (base)
JointSplitModel Dataset 2 split by source: when and how from Dataset 1, who finishes and
who wins on the cards from the round-stats model (the production layer 1s)
The feature lists are the research’s final compact sets (eval/dataset{1,2}_final_features)
minus the market price, which enters as the prior (compact_features).
Research: ufc/modeling/joint.py (models, explained, paired_ci) and
ufc/modeling/cpcv.py (_cells). Log-odds clip 1 % (logit_joint).
Classes
JointRoundModel
X
X(d: pd.DataFrame) -> pd.DataFramechoose_trees
choose_trees(train: pd.DataFrame) -> pd.DataFrameWalk-forward inside train (fit < Y, score Y, Y = 2016-2019), per tree count and
part: mean out-of-sample log loss and its standard error across the folds. The
one-standard-error rule picks the fewest trees within one SE of the best (sets
n_round and n_dec; ko_share_league is left at the last fold’s).
curve
curve(d: pd.DataFrame) -> np.ndarray(rows x 6) S at the end of rounds 0..5: S(0) = 1, S(final bell) = 0 and after.
fit
fit(train: pd.DataFrame) -> JointRoundModelgrid
grid(d: pd.DataFrame) -> dict[str, np.ndarray]For every row (alive at the start of round r): P of every cell from here on. Returns arrays (rows x 5 rounds) for a_ko, a_sub, b_ko, b_sub, and vectors dec, dec_a.
hazard
hazard(d: pd.DataFrame, k: np.ndarray) -> np.ndarrayP(finish in round k+1 | alive at its start) = (1-p0)(S_k - S_k+1) / (p0 + (1-p0) S_k).
prior
prior(d: pd.DataFrame) -> np.ndarrayround_label
round_label(d: pd.DataFrame) -> np.ndarraytargets
targets(d: pd.DataFrame) -> pd.DataFrameJointSplitModel
Bases: JointStatsModel
Dataset 2, split by what each source is good at (validation 2020-2022): whether the round ends and how -> the Dataset 1 round model who is the finisher -> the round-stats model who wins on the cards -> the round-stats decision model
grid
grid(d: pd.DataFrame) -> dict[str, np.ndarray]JointStatsModel
Bases: JointRoundModel
The same grid with round stats (Dataset 2), where they can be known:
current round the round model sees the stats of the rounds already fought later rounds the Dataset 1 round model (their stats cannot be known yet) on the cards P(A wins | decision) sees the stats so far (who is ahead), trained on every round start of fights that went the distance
grid
grid(d: pd.DataFrame) -> dict[str, np.ndarray]Functions
cells
cells(model: JointRoundModel, d: pd.DataFrame) -> pd.DataFrameThe layer-1 columns per row (from here to the end of the fight): the five targets and the six outcome cells a/b x ko/sub/dec.
compact_features
compact_features(final_features: Iterable[str]) -> list[str]The joint model’s features: the research’s final compact set without the market price (the market enters as the prior).
explained
explained(y: Any, p: Any, labels: Sequence[int] = (0, 1)) -> floatShare of log loss explained against the base rate (binary: the mean; multiclass: the class frequencies).
paired_ci
paired_ci(d: pd.DataFrame, model: JointRoundModel, n_boot: int = 2000, seed: int = 0) -> pd.DataFrameJoint model minus benchmark, log loss per row, with a 90% interval from resampling whole cards (events): fights on one card are not independent. Negative = joint better.
softmax
softmax(z: np.ndarray) -> np.ndarray