Skip to content

betting_combat.ml.feature_study

Feature study: which candidate features carry out-of-sample information about a target, and the compact set a model should use. General: a frame, a target, candidate features, a CV splitter and a model; the round system’s wiring is rounds/study.py.

ic_table rank IC (Spearman) of each feature with each target column per period, averaged over the periods; t = mean / sd * sqrt(periods) order_desc a ranking, largest first, NaN last; ties broken by name (a unique key) or, to reproduce the research exactly, by pandas’ default sort redundancy greedy in ranking order: a feature is dropped when its |rank correlation| with an already kept feature exceeds the threshold FeatureStudy out-of-sample scoring on the splitter’s folds: cross_validate log loss (and the training base rate’s), per fold weighted by its test rows; with importance: MDI (gain share) and MDA (log-loss increase when a feature is shuffled within the fold’s test rows) base_logloss the base rate alone (no model fitted) single_feature_importance SFI: every feature on its own, out of sample ablation feature sets in order (e.g. groups added one by one), explained after each importance_table MDI/MDA joined to each feature’s largest IC and t, largest MDA first forward_select candidates in order; one joins when the CV log loss drops by >= min_gain StudyResult one target’s study, saved to / loaded from CSV files

Log loss is scikit-learn’s; explained = 1 - log loss / base-rate log loss. Binary targets are 0/1 columns; a multiclass target names its classes in code order.

Research: ufc/dataset/evaluate.py (FeatureEval: ic_table, select, walk_forward, _permutation, ablation, compact, run).

Classes

CVResult

Out-of-sample log loss (logloss) and the training base rate’s (base_logloss), each the fold log losses weighted by their test rows; folds: per fold (fold, rows, logloss, base_logloss); importance: index feature, gain_share (MDI, mean gain over folds / its sum), perm_logloss_increase (MDA, mean over folds).

FeatureStudy

Out-of-sample feature scoring for one target.

data: the study’s rows (unique index); target: a 0/1 column, or with classes a column of class values (code i = classes[i]); model: fits and predicts; splitter: the folds; perm_seed: MDA shuffles with a fresh default_rng(perm_seed) per fold.

ablation

ablation(steps: Sequence[tuple[Mapping[str, Any], Sequence[str]]], **kwargs: Any) -> pd.DataFrame

Feature sets in order, each (labels, features): one row per step with the labels, logloss, explained and gain_vs_previous (the first step’s is its own).

base_logloss

base_logloss(offset: str | None = None, splitter: Splitter | None = None) -> float

The base rate’s log loss on the folds (as cross_validate’s base_logloss, without fitting a model).

cross_validate

cross_validate(features: Sequence[str], importance: bool = False, offset: str | None = None, params: Mapping[str, Any] | None = None, splitter: Splitter | None = None) -> CVResult

Fit on each fold’s training rows, score its test rows (see CVResult). offset: a log-odds column the model starts from (rows without it are dropped); params override the model’s; splitter overrides the study’s.

encode

encode(s: pd.Series) -> np.ndarray

Target values as integer codes (binary: the 0/1 value itself).

importance_table

importance_table(importance: pd.DataFrame, ic: pd.DataFrame, ties: Ties = 'name', group: bool = True) -> pd.DataFrame

importance (index feature) joined to each feature’s IC and t (over the IC table’s target columns, the value with the largest |.|, each taken on its own) and, with group, its group (ic.group); sorted by MDA, largest first.

rows

rows(offset: str | None = None) -> pd.DataFrame

The study’s rows; with an offset, only those where it is known.

single_feature_importance

single_feature_importance(features: Sequence[str], **kwargs: Any) -> pd.DataFrame

SFI: every feature alone, out of sample (cross_validate keyword arguments pass through). Index feature; columns logloss, explained.

target_columns

target_columns(d: pd.DataFrame | None = None, prefix: str = 'y_') -> pd.DataFrame

The target as numeric columns for the IC: binary as it is (float); multiclass as one 0/1 column per class, named prefix + class.

ForwardResult

The kept features in joining order and the path (feature, logloss, explained, gain: the drop in log loss it brought), from start_logloss.

LightGBMModel

LightGBM classifier. params are the defaults; fit(params=...) overrides them. Binary with an offset: the model starts from the offset (init_score) and P = 1 / (1 + exp(-(offset + raw score))). Multiclass takes no offset.

fit

fit(X: pd.DataFrame, y: np.ndarray, offset: np.ndarray | None = None, params: Mapping[str, Any] | None = None) -> lgb.LGBMClassifier

gain

gain(fitted: lgb.LGBMClassifier) -> np.ndarray

predict

predict(fitted: lgb.LGBMClassifier, X: pd.DataFrame, offset: np.ndarray | None = None) -> np.ndarray

ProbModel

Bases: Protocol

A probability model: fit returns the fitted object; predict gives P(class 1) (binary, shape (n,)) or every class (multiclass, shape (n, K)); gain is the fitted model’s per-feature gain (MDI). offset: a prior on the log-odds scale per row.

fit

fit(X: pd.DataFrame, y: np.ndarray, offset: np.ndarray | None = None, params: Mapping[str, Any] | None = None) -> Any

gain

gain(fitted: Any) -> np.ndarray

predict

predict(fitted: Any, X: pd.DataFrame, offset: np.ndarray | None = None) -> Any

PurgedFolds

PurgedCPCV as a splitter: one fold per combination of test groups (a row is tested in several folds; each fold’s log loss is weighted by its test rows).

folds

folds(d: pd.DataFrame) -> Iterator[Fold]

Splitter

Bases: Protocol

folds

folds(d: pd.DataFrame) -> Iterable[Fold]

StudyResult

One target’s study, as the research wrote it (<tag>_<part>.csv):

ic the IC table (index not written) redundant the redundancy drop log importance MDI / MDA / IC / t / group per feature of the full model (index = feature) ablation the ablation rows compact the forward-selection path summary a one-column table of headline numbers (index = name)

kept: the features after redundancy, in ranking order (not saved; load gives the importance table’s index, in MDA order); compact_kept: the compact set.

load

load(out_dir: Path, tag: str) -> StudyResult

Read a saved study (floats parsed exactly: float_precision='round_trip').

paths

paths(out_dir: Path) -> dict[str, Path]

save

save(out_dir: Path) -> dict[str, Path]

WalkForward

Expanding window by period: for each period P in periods, fit on the rows of earlier periods and score the rows of P. A period without rows is skipped.

folds

folds(d: pd.DataFrame) -> Iterator[Fold]

Functions

forward_select

forward_select(candidates: Sequence[str], score: Callable[[list[str]], CVResult], start_logloss: float, max_features: int, min_gain: float) -> ForwardResult

Candidates in order: one joins when score(kept + [it]) lowers the best log loss so far by at least min_gain; stops at max_features. The path is a frame without columns when nothing joins (as the research wrote it).

ic_table

ic_table(d: pd.DataFrame, features: Sequence[str], Y: pd.DataFrame, period: pd.Series, min_rows: int | None = 50, min_periods: int = 3, drop_missing: bool = True, require_variation: bool = True, finite_only: bool = False) -> pd.DataFrame

Spearman IC of every feature with every column of Y, per period (in sorted order), then the mean over periods and t = mean / sd * sqrt(n).

Per period the rows are the period’s (without the feature’s blanks if drop_missing); a period counts when it has more than min_rows rows (None: no minimum) and, if require_variation, the feature takes more than one value there. finite_only drops periods whose IC is not finite. A feature needs min_periods.

Columns: target (the Y column), feature, ic, t, n_periods, coverage (the feature’s non-blank share over all rows of d). Rows in (target, feature) order.

largest_abs

largest_abs(s: pd.Series) -> Any

The value with the largest absolute value (the first on ties; blanks skipped). All blank: the last value (pandas 2’s s.iloc[s.abs().argmax()], which pandas 3 refuses).

order_desc

order_desc(values: pd.Series, ties: Ties = 'name') -> np.ndarray

Positions of values from largest to smallest, NaN last.

ties='name': equal values in the order of their index labels as text (a unique key, so the order never depends on the sort algorithm). ties='quicksort': pandas’ default sort_values(ascending=False), whose order among equal values is whatever numpy’s quicksort leaves; kept only to reproduce the research outputs exactly.

rank_corr

rank_corr(d: pd.DataFrame, features: Sequence[str]) -> pd.DataFrame

Spearman correlation matrix: Pearson on the ranks, pairwise over non-blank rows.

redundancy

redundancy(order: Sequence[str], corr: pd.DataFrame, score: pd.Series, threshold: float = 0.9) -> tuple[list[str], pd.DataFrame]

Walk order (best first): keep a feature unless its |correlation| with an already kept one exceeds threshold (the first such kept feature, in keeping order, replaces it; a blank correlation never does). Returns the kept list and the drop log (dropped, kept_instead, rho, abs_ic_dropped, abs_ic_kept; score gives the last two). An empty log is a frame without columns (as the research wrote it).