Skip to content

betting_combat.rounds.models.families

Model families for layers 1 and 2 (does another learner beat the boosted-tree correction?). The prior stays in every family: as an offset where the learner supports one (LightGBM init_score), otherwise as input features (the prior’s log-odds, binary, or log-probabilities, multiclass: columns prior_0, prior_1, …). Gaps are filled with the training medians for the learners that need it.

prior only the prior itself lgbm boosted residual, 1 seed; lgbm_bagged: 5 seeds, raw scores averaged logistic L2 (C = 0.1), standardised random_forest / extra_trees: 500 trees, leaves of >= 30 rows, sqrt features

The final model uses ‘lgbm_bagged’ (layer 2’s forward-selected trees) and ‘logistic’ (layer2.layer2_logistic).

Research: ufc/modeling/families.py (fit_predict).

Classes

FittedSklearn

A fitted scikit-learn family (logistic / random forest / extra trees) with the prior as input columns and the training medians that fill gaps: fit_predict split in two, so a model locked on the training rows can price new rows later.

Functions

fit_predict

fit_predict(kind: str, Xtr: pd.DataFrame, ytr: Any, Xte: pd.DataFrame, z_tr: np.ndarray, z_te: np.ndarray, n_classes: int, tree: dict[str, Any]) -> np.ndarray

Fit family kind on the training rows and predict the test rows: P(class 1) (binary) or the class probabilities (multiclass). z_*: the prior as log-odds (binary, 1 column) or log-probabilities (multiclass, K columns); tree: the LightGBM settings.

fit_sklearn

fit_sklearn(kind: str, Xtr: pd.DataFrame, ytr: Any, z_tr: np.ndarray) -> FittedSklearn

Fit family kind (logistic, random_forest, extra_trees) on the training rows.

predict_sklearn

predict_sklearn(fitted: FittedSklearn, Xte: pd.DataFrame, z_te: np.ndarray, n_classes: int) -> np.ndarray

P(class 1) (binary) or the class probabilities (multiclass) for new rows.