datamatters24/ringside-analytics
Ringside Analytics — Pro Wrestling Match Archive A relational snapshot of professional wrestling history from 1980 to the present: 292K matches, 611K wrestler-match participations, 35K events, and 12.8K wrestlers across WWE, AEW, WCW, ECW, NXT, TNA, and others. Sourced from public Cagematch.net scrapes and the alexdiresta profightdb dump, normalized into a Postgres schema, and exported as parquet files that preserve the relational structure (one file per table, joinable by id).… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/ringside-analytics.
Ringside Analytics — Pro Wrestling Match Archive
A relational snapshot of professional wrestling history from 1980 to the present: 292K matches, 611K wrestler-match participations, 35K events, and 12.8K wrestlers across WWE, AEW, WCW, ECW, NXT, TNA, and others. Sourced from public Cagematch.net scrapes and the alexdiresta profightdb dump, normalized into a Postgres schema, and exported as parquet files that preserve the relational structure (one file per table, joinable by id).
This is the source-of-truth companion to the trained model at theodorerubin/ringside-wrestling-archive-match-winner. If you want to train your own model, reshape the features, or just explore 40+ years of booking patterns — start here.
Files
Schema (join keys)
promotions.id ─┬─< wrestlers.primary_promotion_id
├─< events.promotion_id
├─< titles.promotion_id
└─< wrestler_aliases.promotion_id
wrestlers.id ──┬─< match_participants.wrestler_id
├─< wrestler_aliases.wrestler_id
├─< title_reigns.wrestler_id
└─< alignment_turns.wrestler_id
events.id ─────┬─< matches.event_id
└─< alignment_turns.event_id (nullable)
matches.id ────── match_participants.match_id
titles.id ─────── title_reigns.title_idStarter queries
import pandas as pd
matches = pd.read_parquet("matches.parquet")
participants = pd.read_parquet("match_participants.parquet")
wrestlers = pd.read_parquet("wrestlers.parquet")
# Every match The Rock has wrestled, with opponents
rock_id = wrestlers.query("ring_name == 'The Rock'")["id"].iloc[0]
rock_matches = participants[participants["wrestler_id"] == rock_id]-- If you load these into DuckDB:
SELECT w.ring_name, COUNT(*) AS wins
FROM match_participants mp
JOIN wrestlers w ON w.id = mp.wrestler_id
WHERE mp.result = 'win'
GROUP BY 1
ORDER BY 2 DESC
LIMIT 20;Provenance
- Cagematch.net (public HTML scrape, non-commercial use): the bulk of match-level data for 1990-present.
- alexdiresta/all-wwe-and-wwf-matches Kaggle dataset (profightdb dump): cross-validation + pre-1990 coverage.
- Normalization + dedup: entity resolution on wrestler names, match-type classification into a fixed ENUM, and natural-key deduplication to collapse records across sources.
The ETL code and scraper are open source at tedrubin80/wrastlingfirst.
Caveats
- Kayfabe, not athletics. Pro wrestling is scripted. A
resultfield records who was booked to win, not who would win an athletic contest. - Temporal coverage is uneven. 2000-present is well-covered; 1980s are thinner, especially for regional/territory promotions.
- Gender imbalance. Women's division sample size is smaller — expect wider confidence intervals for any women's-division model.
- Ratings are crowd-sourced (Cagematch user ratings). They're a proxy for match quality as perceived by Internet wrestling fans — biased toward work-rate and away from entertainment/story.
License
Released under CC0 1.0 (public domain dedication). Attribution is appreciated but not required. Note that the underlying sources (Cagematch.net, profightdb) have their own terms; this archive is a derivative work made available for research and entertainment.
Citation
@dataset{ringside_analytics_2026,
author = {Rubin, Theodore},
title = {Ringside Analytics: Pro Wrestling Match Archive (1980--present)},
year = {2026},
url = {https://www.kaggle.com/datasets/theodorerubin/ringside-wrestling-archive}
}