betting_combat.ml.feature_study
Feature study: which candidate features carry out-of-sample information about a target,
and the compact set a model should use. General: a frame, a target, candidate features, a CV
splitter and a model; the round system’s wiring is rounds/study.py.
ic_table rank IC (Spearman) of each feature with each target column per period,
averaged over the periods; t = mean / sd * sqrt(periods)
order_desc a ranking, largest first, NaN last; ties broken by name (a unique key)
or, to reproduce the research exactly, by pandas’ default sort
redundancy greedy in ranking order: a feature is dropped when its |rank
correlation| with an already kept feature exceeds the threshold
FeatureStudy out-of-sample scoring on the splitter’s folds:
cross_validate log loss (and the training base rate’s), per fold weighted by its test
rows; with importance: MDI (gain share) and MDA (log-loss increase
when a feature is shuffled within the fold’s test rows)
base_logloss the base rate alone (no model fitted)
single_feature_importance SFI: every feature on its own, out of sample
ablation feature sets in order (e.g. groups added one by one), explained after each
importance_table MDI/MDA joined to each feature’s largest IC and t, largest MDA first
forward_select candidates in order; one joins when the CV log loss drops by >= min_gain
StudyResult one target’s study, saved to / loaded from CSV files
Log loss is scikit-learn’s; explained = 1 - log loss / base-rate log loss. Binary
targets are 0/1 columns; a multiclass target names its classes in code order.
Research: ufc/dataset/evaluate.py (FeatureEval: ic_table, select,
walk_forward, _permutation, ablation, compact, run).
Classes
CVResult
Out-of-sample log loss (logloss) and the training base rate’s
(base_logloss), each the fold log losses weighted by their test rows;
folds: per fold (fold, rows, logloss, base_logloss); importance: index feature,
gain_share (MDI, mean gain over folds / its sum), perm_logloss_increase (MDA, mean over
folds).
FeatureStudy
Out-of-sample feature scoring for one target.
data: the study’s rows (unique index); target: a 0/1 column, or with
classes a column of class values (code i = classes[i]); model: fits and
predicts; splitter: the folds; perm_seed: MDA shuffles with a fresh
default_rng(perm_seed) per fold.
ablation
ablation(steps: Sequence[tuple[Mapping[str, Any], Sequence[str]]], **kwargs: Any) -> pd.DataFrameFeature sets in order, each (labels, features): one row per step with the
labels, logloss, explained and gain_vs_previous (the first step’s is its own).
base_logloss
base_logloss(offset: str | None = None, splitter: Splitter | None = None) -> floatThe base rate’s log loss on the folds (as cross_validate’s base_logloss,
without fitting a model).
cross_validate
cross_validate(features: Sequence[str], importance: bool = False, offset: str | None = None, params: Mapping[str, Any] | None = None, splitter: Splitter | None = None) -> CVResultFit on each fold’s training rows, score its test rows (see CVResult).
offset: a log-odds column the model starts from (rows without it are dropped);
params override the model’s; splitter overrides the study’s.
encode
encode(s: pd.Series) -> np.ndarrayTarget values as integer codes (binary: the 0/1 value itself).
importance_table
importance_table(importance: pd.DataFrame, ic: pd.DataFrame, ties: Ties = 'name', group: bool = True) -> pd.DataFrameimportance (index feature) joined to each feature’s IC and t (over the IC
table’s target columns, the value with the largest |.|, each taken on its own) and,
with group, its group (ic.group); sorted by MDA, largest first.
rows
rows(offset: str | None = None) -> pd.DataFrameThe study’s rows; with an offset, only those where it is known.
single_feature_importance
single_feature_importance(features: Sequence[str], **kwargs: Any) -> pd.DataFrameSFI: every feature alone, out of sample (cross_validate keyword arguments pass
through). Index feature; columns logloss, explained.
target_columns
target_columns(d: pd.DataFrame | None = None, prefix: str = 'y_') -> pd.DataFrameThe target as numeric columns for the IC: binary as it is (float); multiclass as
one 0/1 column per class, named prefix + class.
ForwardResult
The kept features in joining order and the path (feature, logloss, explained, gain:
the drop in log loss it brought), from start_logloss.
LightGBMModel
LightGBM classifier. params are the defaults; fit(params=...) overrides them.
Binary with an offset: the model starts from the offset (init_score) and
P = 1 / (1 + exp(-(offset + raw score))). Multiclass takes no offset.
fit
fit(X: pd.DataFrame, y: np.ndarray, offset: np.ndarray | None = None, params: Mapping[str, Any] | None = None) -> lgb.LGBMClassifiergain
gain(fitted: lgb.LGBMClassifier) -> np.ndarraypredict
predict(fitted: lgb.LGBMClassifier, X: pd.DataFrame, offset: np.ndarray | None = None) -> np.ndarrayProbModel
Bases: Protocol
A probability model: fit returns the fitted object; predict gives P(class 1)
(binary, shape (n,)) or every class (multiclass, shape (n, K)); gain is the fitted
model’s per-feature gain (MDI). offset: a prior on the log-odds scale per row.
fit
fit(X: pd.DataFrame, y: np.ndarray, offset: np.ndarray | None = None, params: Mapping[str, Any] | None = None) -> Anygain
gain(fitted: Any) -> np.ndarraypredict
predict(fitted: Any, X: pd.DataFrame, offset: np.ndarray | None = None) -> AnyPurgedFolds
PurgedCPCV as a splitter: one fold per combination of test groups (a row is
tested in several folds; each fold’s log loss is weighted by its test rows).
folds
folds(d: pd.DataFrame) -> Iterator[Fold]Splitter
Bases: Protocol
folds
folds(d: pd.DataFrame) -> Iterable[Fold]StudyResult
One target’s study, as the research wrote it (<tag>_<part>.csv):
ic the IC table (index not written) redundant the redundancy drop log importance MDI / MDA / IC / t / group per feature of the full model (index = feature) ablation the ablation rows compact the forward-selection path summary a one-column table of headline numbers (index = name)
kept: the features after redundancy, in ranking order (not saved; load gives
the importance table’s index, in MDA order); compact_kept: the compact set.
load
load(out_dir: Path, tag: str) -> StudyResultRead a saved study (floats parsed exactly: float_precision='round_trip').
paths
paths(out_dir: Path) -> dict[str, Path]save
save(out_dir: Path) -> dict[str, Path]WalkForward
Expanding window by period: for each period P in periods, fit on the rows of
earlier periods and score the rows of P. A period without rows is skipped.
folds
folds(d: pd.DataFrame) -> Iterator[Fold]Functions
forward_select
forward_select(candidates: Sequence[str], score: Callable[[list[str]], CVResult], start_logloss: float, max_features: int, min_gain: float) -> ForwardResultCandidates in order: one joins when score(kept + [it]) lowers the best log loss so
far by at least min_gain; stops at max_features. The path is a frame without
columns when nothing joins (as the research wrote it).
ic_table
ic_table(d: pd.DataFrame, features: Sequence[str], Y: pd.DataFrame, period: pd.Series, min_rows: int | None = 50, min_periods: int = 3, drop_missing: bool = True, require_variation: bool = True, finite_only: bool = False) -> pd.DataFrameSpearman IC of every feature with every column of Y, per period (in sorted
order), then the mean over periods and t = mean / sd * sqrt(n).
Per period the rows are the period’s (without the feature’s blanks if
drop_missing); a period counts when it has more than min_rows rows (None: no
minimum) and, if require_variation, the feature takes more than one value there.
finite_only drops periods whose IC is not finite. A feature needs min_periods.
Columns: target (the Y column), feature, ic, t, n_periods, coverage (the feature’s
non-blank share over all rows of d). Rows in (target, feature) order.
largest_abs
largest_abs(s: pd.Series) -> AnyThe value with the largest absolute value (the first on ties; blanks skipped). All
blank: the last value (pandas 2’s s.iloc[s.abs().argmax()], which pandas 3 refuses).
order_desc
order_desc(values: pd.Series, ties: Ties = 'name') -> np.ndarrayPositions of values from largest to smallest, NaN last.
ties='name': equal values in the order of their index labels as text (a unique key,
so the order never depends on the sort algorithm). ties='quicksort': pandas’ default
sort_values(ascending=False), whose order among equal values is whatever numpy’s
quicksort leaves; kept only to reproduce the research outputs exactly.
rank_corr
rank_corr(d: pd.DataFrame, features: Sequence[str]) -> pd.DataFrameSpearman correlation matrix: Pearson on the ranks, pairwise over non-blank rows.
redundancy
redundancy(order: Sequence[str], corr: pd.DataFrame, score: pd.Series, threshold: float = 0.9) -> tuple[list[str], pd.DataFrame]Walk order (best first): keep a feature unless its |correlation| with an already
kept one exceeds threshold (the first such kept feature, in keeping order, replaces
it; a blank correlation never does). Returns the kept list and the drop log
(dropped, kept_instead, rho, abs_ic_dropped, abs_ic_kept; score gives the last two).
An empty log is a frame without columns (as the research wrote it).