Skip to content

betting_combat.consumers.wikipedia

Wikipedia: the ‘List of UFC events’ (every event’s date, page and venue) and each event page’s results table (which fights were main card, prelims, early prelims).

Ported from the research (ufc/pull.py WIKI_URL; ufc/dataset/scrape.py wiki_event_pages, parse_card, _text, SEGMENTS) with identical parsing. Event pages are read through the REST endpoint, one at a time and paced (:class:~.scrape.PacedFetcher, which backs off on 429).

Classes

Wikipedia

card

card(page: str) -> list[dict[str, Any]]

event_list

event_list() -> str

The parse API’s JSON for ‘List of UFC events’, verbatim (validated).

WikipediaFormatError

Bases: ValueError

An answer that is not the page the parsers were written for.

Functions

event_pages

event_pages(list_json: str) -> list[tuple[str, str]]

(date ‘YYYY-MM-DD’, page title) of every event row of the list, in the list’s order, exact duplicates dropped (research wiki_event_pages).

html_text

html_text(fragment: str) -> str

A cell’s text as pandas.read_html (lxml) gives it: every text node (tags and comments dropped, entities decoded), then pandas’ whitespace rule (runs of newlines, or of two or more whitespace characters, become one space; the ends stripped).

list_html

list_html(list_json: str) -> str

The list page’s HTML from the parse API’s JSON; anything else is refused.

list_tables

list_tables(list_json: str) -> list[tuple[list[str], list[list[str]]]]

Every table of the ‘List of UFC events’ page as (header, body rows): the header is the first row, the body the rows after it (research pull_events read them with pandas.read_html: the scheduled events have ‘Event’ and ‘Date’ columns, the past events ’#’, ‘Event’ and ‘Date’).

page_url

page_url(page: str) -> str

The REST address of one event page (the research quoted the title with no safe characters).

parse_card

parse_card(page_html: str) -> list[dict[str, Any]]

Rows of the results table: segment header rows switch the segment; fight rows are ’ | | def./vs. | | | |

parse_card_bouts

parse_card_bouts(page_html: str) -> list[dict[str, Any]]

parse_card’s rows with the weight-class cell and the champion marks kept: each fight row of the results (or scheduled fight card) table as segment, card_order (1 = the main event), weight_class (Wikipedia’s text, e.g. “Women’s Strawweight”), fighter_1 / fighter_2 (the ‘(c)’ / ‘(ic)’ marks removed), champion (either name carried the mark) and notes (the row’s cells after the second fighter, joined: method, round, time and notes such as ‘For the UFC Lightweight Championship’).

table_grid

table_grid(table_html: str) -> list[list[str]]

A table’s rows as cell texts (html_text), a cell spanning rows or columns repeated in every position it covers (as pandas.read_html fills them).