Skip to content

betting_combat.consumers.scrape

Polite GETs for the public pages the round dataset scrapes (Wikipedia, MMADecisions).

One request at a time, at most one every pause_s seconds (the research paced ~1.5 s); 429 and 5xx are retried with backoff, honouring Retry-After (Wikipedia has answered 429 before); every other HTTP error fails at once. URLs are quoted exactly as the research quoted them (unicode fighter names in MMADecisions slugs) and bodies decoded as UTF-8 with replacement, as the research decoded them.

Classes

PacedFetcher

text

text(url: str) -> str

ScrapeError

Bases: RuntimeError

Functions

quote_url

quote_url(url: str) -> str

The research’s quoting: every character outside :/?&=%#+ and the unreserved set is percent-encoded (UTF-8), existing escapes are kept.

retry_after_seconds

retry_after_seconds(value: str | None, now: dt.datetime | None = None) -> float | None

A Retry-After header as seconds to wait: delay-seconds, or an HTTP-date (the time left until it, never negative). None when absent or unreadable.