betting_combat.services.rounds_datasets
Strategy 4’s model-ready datasets: built from the raw tables by the datasets job
(jobs.rounds.DatasetsFlow) into three dataset tables of betting, and read back by the
refit, which trains on them and on nothing else.
rounds_datasets the whole-history products of ONE build, replaced together in one transaction: Dataset 1 and 2 (Groups 1-5 and the targets; Group 6, an evaluation-only input, is left out), the feature catalogue, Group 5's model features, the Kalshi fight clock and its market list, Set 4 (the minute panel), Set 3 main and side (keyed, the spreads filled over the whole build, the unpriceable side rows dropped), the tape-fill grid and each fight's last trade read (the refit's cutoff guard). One row per product: the frame pickled (zlib, base64; its SHA-256 checked on every read) and the build's manifest.rounds_fight_frames per Kalshi-listed fight: its Set 4, its Set 3 rows before the spread fill and its fills, keyed by the digest of exactly its inputs (``services.rounds_cache``): the per-fight partitions a build reuses, so an update rebuilds only the fights whose inputs changedrounds_clock_years per year: Group 5's clock fit (the p0 model and the Bayesian finish curve on the bouts before 1 January, seed = the year), keyed by the digest of its inputs: fitted once, and the refit's artifact carries the card year'sA build is services.rounds_frames.build_frames (the research build order, every step the
ported library’s), unchanged: these tables hold its output, so the datasets are the research’s
(tests/store/test_rounds_datasets_known_answers.py).
The manifest says what the build read: its basis is the digest of every raw surface’s size
and latest time (RoundsStore.surface_counts), the mirror’s, the static files’ and the
Wikipedia list’s SHA-256, the source of the code that builds the datasets and the library
versions. The refit refuses datasets whose basis is not the tables’ current one (the raw tables
changed, or the code did, since the build) and a card year whose clock the build did not fit.
Classes
DatasetsReport
DatasetsUnavailable
Bases: RuntimeError
The dataset tables cannot serve a refit: none built, a product missing or damaged, rows of two builds, or no clock fit for the year asked.
Manifest
What one build read and produced; stored on every product row.
from_json
from_json(raw: Mapping[str, Any]) -> Manifestto_json
to_json() -> dict[str, Any]RoundsDatasetsService
One build of the dataset tables (see the module doc).
build
build(reuse: bool = True, model_features: pd.DataFrame | None = None, card: dt.date | None = None) -> DatasetsReportBuild every product from the raw tables and replace rounds_datasets with them.
reuse: read the stored per-fight frames and clock fits whose inputs are unchanged
(an update); False rebuilds every fight and refits every year (a build), and stores
them. model_features: Group 5 given whole (the research’s; the known-answer tests),
instead of the clock fits. card: the card the refit after it locks to (a prep
chain’s; None: the next card in the stored list): its year’s clock is fitted. Run it under the data lock: the basis is read first, then
the tables, and nothing may land in between.
Functions
load_datasets
load_datasets(store: RoundsStore, clock_years: Sequence[int]) -> tuple[Manifest, Frames]The stored build as Frames (each product’s SHA-256 checked), with the clock fits of
clock_years (each the very fit the build priced with, fitted under the running
libraries). DatasetsUnavailable when it cannot (see the class).
load_manifest
load_manifest(store: RoundsStore) -> Manifest | NoneThe stored build’s manifest (None: no build stored); DatasetsUnavailable when the
rows are not one whole build.
product_rows
product_rows(frames: Frames, manifest: Manifest) -> list[dict[str, Any]]One rounds_datasets row per product of frames.
raw_basis
raw_basis(store: RoundsStore) -> strThe digest of what a build reads (see the module doc).