betting_combat.rounds.study
The round datasets’ feature study: which features carry information about which target, measured on the training years only, and the compact feature set each target keeps.
Time split: train 2008-2019 | validation 2020-2022 | test 2023+. Everything here reads the training years only.
Per target (RoundFeatureStudy.run_target):
- IC Spearman IC of each candidate with the target per year, averaged; t
- redundancy candidates by |IC| (best over the target’s columns); one correlated |rho| > 0.9 (ranks, all training-year rows) with a better one is dropped
- LightGBM walk-forward inside train (fit on years < Y, score Y, Y = 2014..2019): out-of- sample log loss vs the base rate, gain share (MDI), permutation importance (MDA)
- ablation groups added in order (context -> odds -> fighter -> matchup -> model -> round stats): out-of-sample log loss after each
- benchmark the market price (winner) or the bookmakers’ method odds (method)
- compact forward selection over the candidates with MDA > 0 and |IC t| >= 2, in MDA
order, walk-forward on 2016-2019 with a smaller model: a feature joins when
it lowers the log loss by >= 0.0005 per row. The winner starts from the
market’s logit (offset), so only what beats the price survives.
Then
final_feature_set: the union of every target’s compact set (the research froze it aseval/dataset{1,2}_final_features.csv, the joint model’s features).
Target rows: y_win_fight at round_idx 0 (Dataset 2: every round start), y_method at
round_idx 0 (three classes), y_finish_round / y_decision / y_win_round_p at every round start.
Candidates: the catalogue’s numeric columns except keys, targets and live prices (Group 6);
symmetric only (no _a / _b column except the market price), round_idx first.
Also: fighter_study (the fighter panel on its own), target_balance, the report’s
matrices (compact_union, ic_matrix, compact_corr), and layer 2’s own study on its
purged CPCV (layer2_study, layer2_value_test, layer2_run: the non-nested
selection and the value of round stats / side markets; models/layer2.nested is the
honest estimate).
Ties: ties='name' (default) breaks equal |IC| / equal MDA by feature name;
ties='quicksort' reproduces the research exactly (pandas’ default sort, whose order among
equal values is numpy’s quicksort’s; e.g. rounds_left and five_rounds have equal |IC| for the
winner and the method at round 0).
Research: ufc/dataset/evaluate.py (FeatureEval, target_balance,
evaluate_fighters, final_feature_set) and ufc/dataset/report.py
(matrices_section: the computations only), ufc/modeling/layer2.py (Layer2.study,
Layer2.value_test, main).
Classes
RoundFeatureStudy
The study of one round dataset (name: ‘dataset1’ or ‘dataset2’) against its
catalogue (column, group, known).
ablation
ablation(target: str) -> pd.DataFramebenchmark
benchmark(target: str) -> dict[str, Any]The best simple reference on the same CV rows: the market’s own price for the winner (rows with a price), the bookmakers’ method prices for the method (rows with them). None where no market exists historically.
compact
compact(target: str, imp: pd.DataFrame, max_features: int = COMPACT_MAX_FEATURES, min_gain: float = COMPACT_MIN_GAIN) -> dict[str, Any]Forward selection over the candidates (MDA > 0 and |IC t| >= 2), in the
importance table’s order: a feature joins only if it lowers the walk-forward
(2016-2019) log loss by at least min_gain per row. For the winner the model
starts from the market’s logit, so only what beats the price survives.
Returns kept, path, explained (the kept set walk-forward on 2014-2019; nothing kept: round_idx alone, or for the winner the market price alone), offset, start_logloss.
features
features(groups: Sequence[str] | None = None) -> list[str]Candidates in catalogue order (round_idx first), optionally of some groups only.
ic_table
ic_table(target: str) -> pd.DataFrameColumns target, feature, ic, t, years, coverage, group, abs_ic; by target, then |IC| largest first (equal |IC| in candidate order: the sort is stable).
run
run(targets: Sequence[str] = tuple(TARGETS), out_dir: Path | None = None) -> dict[str, StudyResult]Every target’s study; with out_dir each is saved as it finishes.
run_target
run_target(target: str) -> StudyResultselect
select(ic: pd.DataFrame, threshold: float = REDUNDANCY_THRESHOLD) -> tuple[list[str], pd.DataFrame]Greedy by |IC| (best over the target’s columns): keep a feature unless it is
correlated beyond threshold with one already kept (ranks, every training-year
row). Returns the kept features and the drop log.
study
study(target: str) -> FeatureStudyThe walk-forward LightGBM study of target on its training rows.
train_rows
train_rows(target: str) -> pd.DataFrameFunctions
compact_corr
compact_corr(d: pd.DataFrame, feats: Sequence[str]) -> pd.DataFrameSpearman correlation of feats on the training years.
compact_union
compact_union(results: Mapping[str, StudyResult]) -> list[str]Every target’s compact features, in target order, each once.
fighter_study
fighter_study(fp: pd.DataFrame, blocks: Mapping[str, Sequence[str]], targets: Sequence[str] = FIGHTER_TARGETS, ties: Ties = 'name') -> dict[str, dict[str, Any]]The fighter on his own (no opponent), training years only: per fighter-level target,
the IC per year (years with > 50 rows, finite ICs only, at least 3) and the walk-forward
LightGBM on every history feature (blocks, those not all blank) with permutation
importance. Per target: explained and importance (index feature: ic, t,
perm_logloss_increase; features with an IC only, largest MDA first).
final_feature_set
final_feature_set(out_dir: Path, name: str, catalogue: pd.DataFrame, write: bool = True) -> pd.DataFramefinal_feature_table of the saved compact paths, written to
<name>_final_features.csv (as the research did).
final_feature_table
final_feature_table(compacts: Mapping[str, pd.DataFrame], catalogue: pd.DataFrame, targets: Sequence[str] | None = None) -> pd.DataFrameThe union of every target’s compact set (compacts: target -> compact path):
one row per feature (index feature) with its group, when it is known, its
out-of-sample gain for each of targets (default: the compacts’ keys; a column per
target, in that order) that chose it (log loss per row x 1e3; blank: not chosen) and
n_targets; by group, then n_targets largest first.
ic_matrix
ic_matrix(results: Mapping[str, StudyResult], feats: Sequence[str]) -> pd.DataFrameThe IC of feats (rows) against every target column (columns: the IC tables’
target names, per target in sorted order, e.g. y_dec, y_ko, y_sub for the method).
layer2_run
layer2_run(L: Layer2, ties: Ties = 'name') -> dict[str, Any]Every layer-2 target’s study and value tests, as the research’s main wrote them:
per target importance / forward / ablation / oos, and the summary and value-test
tables.
layer2_study
layer2_study(L: Layer2, target: str, ties: Ties = 'name') -> dict[str, Any]Layer 2’s (non-nested) study of target on its CPCV: IC of the core candidates
(l1, trade, set1; the prior itself left out), redundancy by |IC|, MDA of the survivors,
forward selection over MDA > 0 and |t| >= 1.5 (MDA order), the kept set’s group
ablation (l1 -> + trade -> + set1) and final vs prior with a 90 % card interval.
Selection bias: the selection and its score share the folds (nested is the honest
estimate). Returns ic, importance (mda, ic, t, group), kept, path, ablation,
prior_logloss, final_logloss, final_vs_prior (estimate, low, high, cards), p_final,
p_prior.
layer2_value_test
layer2_value_test(L: Layer2, target: str, kept: list[str], extra: str) -> dict[str, Any]The core compact set vs core + a whole extra group (‘stats’ or ‘side’; ‘side’ on the rows with side markets): the extra group’s candidates with |t| >= 1.5 in |IC| order (the first 25) forward-selected on the same CV, then the difference with a card interval.
read_compacts
read_compacts(out_dir: Path, name: str, targets: Sequence[str] = tuple(TARGETS)) -> dict[str, pd.DataFrame]The saved compact paths of every target that has one (an empty file = none), read with pandas’ default float parser as the research read them (it can be one unit in the last place off the written value; the frozen final feature sets carry that).
target_balance
target_balance(d: pd.DataFrame) -> pd.DataFramePer target: its rows (round 0 or every round start) and the share of each value.
year_of
year_of(d: pd.DataFrame) -> pd.Series