betting_combat.ml.lopez
Sample weights and bootstrap for overlapping labels (Lopez de Prado, AFML ch. 4).
A bet is labelled by an outcome that “lives” from its entry until the market settles. Bets on the same fight overlap in time and share their outcome; bets on different fights never share a label, so concurrency is counted per fight, minute by minute.
indicator (events x fight-minutes) 1 where the bet is open average_uniqueness for each bet, the mean over its lifetime of 1 / (bets of that fight open at that minute) (AFML 4.4) sequential_bootstrap draws rows one at a time with probability proportional to their average uniqueness given the rows already drawn (AFML 4.5) SeqBootBagger a bagged ensemble whose bags come from the sequential bootstrap
Research: ufc/modeling/lopez.py.
Classes
SeqBootBagger
Bases: BaseEstimator, ClassifierMixin
Bagging with sequential-bootstrap bags (AFML 4.5 + ch. 6): fit(X, y, ind=...)
with ind the training rows’ indicator matrix.
fit
fit(X: Any, y: Any, ind: sparse.csr_matrix, sample_weight: Any = None) -> SeqBootBaggerpredict_proba
predict_proba(X: Any) -> np.ndarrayFunctions
average_uniqueness
average_uniqueness(ind: sparse.csr_matrix) -> np.ndarrayindicator
indicator(ev: pd.DataFrame, start: str = 't_entry', end: str = 't_end') -> sparse.csr_matrix(events x fight-minutes) 1 where the bet is open. Columns are per fight (in order of first appearance), so bets on different fights never overlap.
sequential_bootstrap
sequential_bootstrap(ind: sparse.csr_matrix, n: int | None = None, rng: np.random.Generator | None = None) -> np.ndarray