Skip to content

betting_combat.rounds.study

The round datasets’ feature study: which features carry information about which target, measured on the training years only, and the compact feature set each target keeps.

Time split: train 2008-2019 | validation 2020-2022 | test 2023+. Everything here reads the training years only.

Per target (RoundFeatureStudy.run_target):

  1. IC Spearman IC of each candidate with the target per year, averaged; t
  2. redundancy candidates by |IC| (best over the target’s columns); one correlated |rho| > 0.9 (ranks, all training-year rows) with a better one is dropped
  3. LightGBM walk-forward inside train (fit on years < Y, score Y, Y = 2014..2019): out-of- sample log loss vs the base rate, gain share (MDI), permutation importance (MDA)
  4. ablation groups added in order (context -> odds -> fighter -> matchup -> model -> round stats): out-of-sample log loss after each
  5. benchmark the market price (winner) or the bookmakers’ method odds (method)
  6. compact forward selection over the candidates with MDA > 0 and |IC t| >= 2, in MDA order, walk-forward on 2016-2019 with a smaller model: a feature joins when it lowers the log loss by >= 0.0005 per row. The winner starts from the market’s logit (offset), so only what beats the price survives. Then final_feature_set: the union of every target’s compact set (the research froze it as eval/dataset{1,2}_final_features.csv, the joint model’s features).

Target rows: y_win_fight at round_idx 0 (Dataset 2: every round start), y_method at round_idx 0 (three classes), y_finish_round / y_decision / y_win_round_p at every round start. Candidates: the catalogue’s numeric columns except keys, targets and live prices (Group 6); symmetric only (no _a / _b column except the market price), round_idx first.

Also: fighter_study (the fighter panel on its own), target_balance, the report’s matrices (compact_union, ic_matrix, compact_corr), and layer 2’s own study on its purged CPCV (layer2_study, layer2_value_test, layer2_run: the non-nested selection and the value of round stats / side markets; models/layer2.nested is the honest estimate).

Ties: ties='name' (default) breaks equal |IC| / equal MDA by feature name; ties='quicksort' reproduces the research exactly (pandas’ default sort, whose order among equal values is numpy’s quicksort’s; e.g. rounds_left and five_rounds have equal |IC| for the winner and the method at round 0).

Research: ufc/dataset/evaluate.py (FeatureEval, target_balance, evaluate_fighters, final_feature_set) and ufc/dataset/report.py (matrices_section: the computations only), ufc/modeling/layer2.py (Layer2.study, Layer2.value_test, main).

Classes

RoundFeatureStudy

The study of one round dataset (name: ‘dataset1’ or ‘dataset2’) against its catalogue (column, group, known).

ablation

ablation(target: str) -> pd.DataFrame

benchmark

benchmark(target: str) -> dict[str, Any]

The best simple reference on the same CV rows: the market’s own price for the winner (rows with a price), the bookmakers’ method prices for the method (rows with them). None where no market exists historically.

compact

compact(target: str, imp: pd.DataFrame, max_features: int = COMPACT_MAX_FEATURES, min_gain: float = COMPACT_MIN_GAIN) -> dict[str, Any]

Forward selection over the candidates (MDA > 0 and |IC t| >= 2), in the importance table’s order: a feature joins only if it lowers the walk-forward (2016-2019) log loss by at least min_gain per row. For the winner the model starts from the market’s logit, so only what beats the price survives.

Returns kept, path, explained (the kept set walk-forward on 2014-2019; nothing kept: round_idx alone, or for the winner the market price alone), offset, start_logloss.

features

features(groups: Sequence[str] | None = None) -> list[str]

Candidates in catalogue order (round_idx first), optionally of some groups only.

ic_table

ic_table(target: str) -> pd.DataFrame

Columns target, feature, ic, t, years, coverage, group, abs_ic; by target, then |IC| largest first (equal |IC| in candidate order: the sort is stable).

run

run(targets: Sequence[str] = tuple(TARGETS), out_dir: Path | None = None) -> dict[str, StudyResult]

Every target’s study; with out_dir each is saved as it finishes.

run_target

run_target(target: str) -> StudyResult

select

select(ic: pd.DataFrame, threshold: float = REDUNDANCY_THRESHOLD) -> tuple[list[str], pd.DataFrame]

Greedy by |IC| (best over the target’s columns): keep a feature unless it is correlated beyond threshold with one already kept (ranks, every training-year row). Returns the kept features and the drop log.

study

study(target: str) -> FeatureStudy

The walk-forward LightGBM study of target on its training rows.

train_rows

train_rows(target: str) -> pd.DataFrame

Functions

compact_corr

compact_corr(d: pd.DataFrame, feats: Sequence[str]) -> pd.DataFrame

Spearman correlation of feats on the training years.

compact_union

compact_union(results: Mapping[str, StudyResult]) -> list[str]

Every target’s compact features, in target order, each once.

fighter_study

fighter_study(fp: pd.DataFrame, blocks: Mapping[str, Sequence[str]], targets: Sequence[str] = FIGHTER_TARGETS, ties: Ties = 'name') -> dict[str, dict[str, Any]]

The fighter on his own (no opponent), training years only: per fighter-level target, the IC per year (years with > 50 rows, finite ICs only, at least 3) and the walk-forward LightGBM on every history feature (blocks, those not all blank) with permutation importance. Per target: explained and importance (index feature: ic, t, perm_logloss_increase; features with an IC only, largest MDA first).

final_feature_set

final_feature_set(out_dir: Path, name: str, catalogue: pd.DataFrame, write: bool = True) -> pd.DataFrame

final_feature_table of the saved compact paths, written to <name>_final_features.csv (as the research did).

final_feature_table

final_feature_table(compacts: Mapping[str, pd.DataFrame], catalogue: pd.DataFrame, targets: Sequence[str] | None = None) -> pd.DataFrame

The union of every target’s compact set (compacts: target -> compact path): one row per feature (index feature) with its group, when it is known, its out-of-sample gain for each of targets (default: the compacts’ keys; a column per target, in that order) that chose it (log loss per row x 1e3; blank: not chosen) and n_targets; by group, then n_targets largest first.

ic_matrix

ic_matrix(results: Mapping[str, StudyResult], feats: Sequence[str]) -> pd.DataFrame

The IC of feats (rows) against every target column (columns: the IC tables’ target names, per target in sorted order, e.g. y_dec, y_ko, y_sub for the method).

layer2_run

layer2_run(L: Layer2, ties: Ties = 'name') -> dict[str, Any]

Every layer-2 target’s study and value tests, as the research’s main wrote them: per target importance / forward / ablation / oos, and the summary and value-test tables.

layer2_study

layer2_study(L: Layer2, target: str, ties: Ties = 'name') -> dict[str, Any]

Layer 2’s (non-nested) study of target on its CPCV: IC of the core candidates (l1, trade, set1; the prior itself left out), redundancy by |IC|, MDA of the survivors, forward selection over MDA > 0 and |t| >= 1.5 (MDA order), the kept set’s group ablation (l1 -> + trade -> + set1) and final vs prior with a 90 % card interval.

Selection bias: the selection and its score share the folds (nested is the honest estimate). Returns ic, importance (mda, ic, t, group), kept, path, ablation, prior_logloss, final_logloss, final_vs_prior (estimate, low, high, cards), p_final, p_prior.

layer2_value_test

layer2_value_test(L: Layer2, target: str, kept: list[str], extra: str) -> dict[str, Any]

The core compact set vs core + a whole extra group (‘stats’ or ‘side’; ‘side’ on the rows with side markets): the extra group’s candidates with |t| >= 1.5 in |IC| order (the first 25) forward-selected on the same CV, then the difference with a card interval.

read_compacts

read_compacts(out_dir: Path, name: str, targets: Sequence[str] = tuple(TARGETS)) -> dict[str, pd.DataFrame]

The saved compact paths of every target that has one (an empty file = none), read with pandas’ default float parser as the research read them (it can be one unit in the last place off the written value; the frozen final feature sets carry that).

target_balance

target_balance(d: pd.DataFrame) -> pd.DataFrame

Per target: its rows (round 0 or every round start) and the share of each value.

year_of

year_of(d: pd.DataFrame) -> pd.Series