Skip to content

betting_combat.rounds.features.similarity

Similarity: fighters grouped by style, and how styles have fared against each other.

Port of the research ufc/dataset/similarity.py; the numbers must stay identical.

style clusters k-means (k = 6) on each fighter’s as-of style vector: strikes landed and absorbed per minute, striking accuracy, takedowns per 15, takedown defence, submission attempts per 15, control share, knockdowns per 15. Fitted on fighter-fights before 2020 with at least 3 earlier fights (train years only); every fighter-fight is then assigned to its nearest centre. style_edge_a how fighters of A’s style have done against fighters of B’s style in earlier fights: win rate minus 0.5, shrunk toward 0 with 20 pseudo-fights (0 when both share a style, and for A’s first fights) style_finish how often fights between those two styles ended inside the distance, from earlier fights, shrunk toward the league rate with 20 pseudo-fights style_known 1 when both fighters had 2+ earlier UFC fights (a style estimate means something); the tree can discount the rest

A new fighter with a known style gets the record of fighters who fought like him: the cross-sectional learning a raw fighter id cannot give.

Classes

StyleSimilarity

build

build() -> pd.DataFrame

Per fight_id: style_1, style_2, style_edge_1 (from slot 1’s side), style_finish, style_known.

centres

centres() -> pd.DataFrame

The six styles in the original units, for reading them.

clusters

clusters() -> pd.DataFrame