betting_combat.rounds.models.families
Model families for layers 1 and 2 (does another learner beat the boosted-tree
correction?). The prior stays in every family: as an offset where the learner supports one
(LightGBM init_score), otherwise as input features (the prior’s log-odds, binary, or
log-probabilities, multiclass: columns prior_0, prior_1, …). Gaps are filled with
the training medians for the learners that need it.
prior only the prior itself lgbm boosted residual, 1 seed; lgbm_bagged: 5 seeds, raw scores averaged logistic L2 (C = 0.1), standardised random_forest / extra_trees: 500 trees, leaves of >= 30 rows, sqrt features
The final model uses ‘lgbm_bagged’ (layer 2’s forward-selected trees) and ‘logistic’
(layer2.layer2_logistic).
Research: ufc/modeling/families.py (fit_predict).
Classes
FittedSklearn
A fitted scikit-learn family (logistic / random forest / extra trees) with the prior
as input columns and the training medians that fill gaps: fit_predict split in two,
so a model locked on the training rows can price new rows later.
Functions
fit_predict
fit_predict(kind: str, Xtr: pd.DataFrame, ytr: Any, Xte: pd.DataFrame, z_tr: np.ndarray, z_te: np.ndarray, n_classes: int, tree: dict[str, Any]) -> np.ndarrayFit family kind on the training rows and predict the test rows: P(class 1)
(binary) or the class probabilities (multiclass). z_*: the prior as log-odds (binary,
1 column) or log-probabilities (multiclass, K columns); tree: the LightGBM settings.
fit_sklearn
fit_sklearn(kind: str, Xtr: pd.DataFrame, ytr: Any, z_tr: np.ndarray) -> FittedSklearnFit family kind (logistic, random_forest, extra_trees) on the training rows.
predict_sklearn
predict_sklearn(fitted: FittedSklearn, Xte: pd.DataFrame, z_te: np.ndarray, n_classes: int) -> np.ndarrayP(class 1) (binary) or the class probabilities (multiclass) for new rows.