betting_combat.consumers.wikipedia
Wikipedia: the ‘List of UFC events’ (every event’s date, page and venue) and each event page’s results table (which fights were main card, prelims, early prelims).
Ported from the research (ufc/pull.py WIKI_URL; ufc/dataset/scrape.py
wiki_event_pages, parse_card, _text, SEGMENTS) with identical parsing.
Event pages are read through the REST endpoint, one at a time and paced
(:class:~.scrape.PacedFetcher, which backs off on 429).
Classes
Wikipedia
card
card(page: str) -> list[dict[str, Any]]event_list
event_list() -> strThe parse API’s JSON for ‘List of UFC events’, verbatim (validated).
WikipediaFormatError
Bases: ValueError
An answer that is not the page the parsers were written for.
Functions
event_pages
event_pages(list_json: str) -> list[tuple[str, str]](date ‘YYYY-MM-DD’, page title) of every event row of the list, in the list’s
order, exact duplicates dropped (research wiki_event_pages).
html_text
html_text(fragment: str) -> strA cell’s text as pandas.read_html (lxml) gives it: every text node (tags and
comments dropped, entities decoded), then pandas’ whitespace rule (runs of newlines, or of
two or more whitespace characters, become one space; the ends stripped).
list_html
list_html(list_json: str) -> strThe list page’s HTML from the parse API’s JSON; anything else is refused.
list_tables
list_tables(list_json: str) -> list[tuple[list[str], list[list[str]]]]Every table of the ‘List of UFC events’ page as (header, body rows): the header is
the first row, the body the rows after it (research pull_events read them with
pandas.read_html: the scheduled events have ‘Event’ and ‘Date’ columns, the past
events ’#’, ‘Event’ and ‘Date’).
page_url
page_url(page: str) -> strThe REST address of one event page (the research quoted the title with no safe characters).
parse_card
parse_card(page_html: str) -> list[dict[str, Any]]Rows of the results table: segment header rows switch the segment; fight rows are
’
parse_card_bouts
parse_card_bouts(page_html: str) -> list[dict[str, Any]]parse_card’s rows with the weight-class cell and the champion marks kept: each
fight row of the results (or scheduled fight card) table as segment, card_order
(1 = the main event), weight_class (Wikipedia’s text, e.g. “Women’s Strawweight”),
fighter_1 / fighter_2 (the ‘(c)’ / ‘(ic)’ marks removed), champion (either
name carried the mark) and notes (the row’s cells after the second fighter, joined:
method, round, time and notes such as ‘For the UFC Lightweight Championship’).
table_grid
table_grid(table_html: str) -> list[list[str]]A table’s rows as cell texts (html_text), a cell spanning rows or columns repeated
in every position it covers (as pandas.read_html fills them).