Skip to content

betting_combat.ml.lopez

Sample weights and bootstrap for overlapping labels (Lopez de Prado, AFML ch. 4).

A bet is labelled by an outcome that “lives” from its entry until the market settles. Bets on the same fight overlap in time and share their outcome; bets on different fights never share a label, so concurrency is counted per fight, minute by minute.

indicator (events x fight-minutes) 1 where the bet is open average_uniqueness for each bet, the mean over its lifetime of 1 / (bets of that fight open at that minute) (AFML 4.4) sequential_bootstrap draws rows one at a time with probability proportional to their average uniqueness given the rows already drawn (AFML 4.5) SeqBootBagger a bagged ensemble whose bags come from the sequential bootstrap

Research: ufc/modeling/lopez.py.

Classes

SeqBootBagger

Bases: BaseEstimator, ClassifierMixin

Bagging with sequential-bootstrap bags (AFML 4.5 + ch. 6): fit(X, y, ind=...) with ind the training rows’ indicator matrix.

fit

fit(X: Any, y: Any, ind: sparse.csr_matrix, sample_weight: Any = None) -> SeqBootBagger

predict_proba

predict_proba(X: Any) -> np.ndarray

Functions

average_uniqueness

average_uniqueness(ind: sparse.csr_matrix) -> np.ndarray

indicator

indicator(ev: pd.DataFrame, start: str = 't_entry', end: str = 't_end') -> sparse.csr_matrix

(events x fight-minutes) 1 where the bet is open. Columns are per fight (in order of first appearance), so bets on different fights never overlap.

sequential_bootstrap

sequential_bootstrap(ind: sparse.csr_matrix, n: int | None = None, rng: np.random.Generator | None = None) -> np.ndarray