Team Ai
Datasetpublic

dingjiacheng/wc2026-agents

WC2026-Agents: ChatGPT vs Claude vs Gemini vs Grok on the 2026 FIFA World Cup WC2026-Agents is a contamination-free benchmark in which four frontier LLMs act as autonomous forecasting agents over the entire 2026 FIFA World Cup (104 matches, 11 June to 19 July 2026). Each agent (Claude Opus 4.8, ChatGPT GPT-5.5 with high reasoning, Gemini 3.1 Pro, and Grok Expert Mode) ran an identical search, act, reflect loop per match: search the web, commit to a 1X2 (team A win / draw / team… See the full description on the dataset page: https://huggingface.co/datasets/dingjiacheng/wc2026-agents.

sourceHugging Facecc-by-4.0updated 26d agoView on Hugging Face
0likes76downloads
Dataset Card

WC2026-Agents: ChatGPT vs Claude vs Gemini vs Grok on the 2026 FIFA World Cup

WC2026-Agents is a contamination-free benchmark in which four frontier LLMs act as autonomous forecasting agents over the entire 2026 FIFA World Cup (104 matches, 11 June to 19 July 2026). Each agent (Claude Opus 4.8, ChatGPT GPT-5.5 with high reasoning, Gemini 3.1 Pro, and Grok Expert Mode) ran an identical search, act, reflect loop per match: search the web, commit to a 1X2 (team A win / draw / team B win) probability and a virtual bet of up to $100, and after the match reflect given only the final score. Every match kicked off after the models' training cutoffs. The pre-match betting market is included as a fifth forecaster via per-match 1X2 odds.

πŸ“„ Paper: arXiv:2607.17765, FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches, Jiacheng Ding and Cong Guo (University of Memphis) 🌐 Project page: https://graphuofm.github.io/FIFA2026LLM/ (δΈ­ζ–‡) πŸ’» Code: https://github.com/graphuofm/FIFA2026LLM

Headline results

ForecasterAccuracyBrier (lower is better)Log-lossVirtual profitROIBets
Betting market (vig-free odds)68.3%0.4690.807+$1,041\*104
Grok (Expert Mode)68.3%0.47060.803+$650+10.3%104
Gemini 3.1 Pro65.4%0.48280.820+$322+3.7%103
ChatGPT (GPT-5.5)68.3%0.47290.814+$118+8.0%55
Claude Opus 4.866.3%0.47050.806-$275-18.1%73

\*Flat stake on the market favourite in every match.

  • β€”The four agents pick the same outcome in 96 of 104 matches (92%); none beats the market's Brier score, and a flat bet on the market favourite out-earns all four.
  • β€”They diverge as decision-makers: betting ROI from -18.1% to +10.3%. Bets against the market favourite win 21% to 40% of the time (vs 48% to 69% with the market); they lose money for Claude, Gemini and ChatGPT, while Grok's 5 contrarian bets make +$64.
  • β€”Market odds are cited in 100% of Claude's forecasts but only 12% of Gemini's.
  • β€”Draws: 24 of 104 matches were 90-minute draws, but no agent made a draw its top pick more than 4 times.
  • β€”Self-reflection: on wrong picks, Gemini admits "incorrect" 86% of the time, ChatGPT 36%.

What's inside

ConfigFileRowsWhat
forecasts (default)data/forecasts.csv416One row per (agent, match): probabilities, bet, verbatim reasoning, the reflection fields, outcome, market-implied probs, and settled P&L
metricsdata/model_metrics.csv5Accuracy / Brier / log-loss / ECE per agent + market
betting_summarydata/betting_summary.csv12ROI, profit, hit rate per agent and stage
final_factor_probedata/final_factor_probe.csv7Final/bronze-only "which narrative factors influenced you" probe
schedulemetadata/schedule.csv104Fixtures (group + knockout), stage, venue, kickoff
resultsmetadata/results.csv104Ground truth incl. penalty shootouts and who advanced
oddsmetadata/odds.csv104Pre-match 1X2 odds with source URL, vig-removed implied probs
raw/raw/*.txtOriginal per-model transcripts (read-only) + the prompt templates

Coverage: 4 agents Γ— 104 matches = 416 forecasts and 414 reflections (two Gemini group reflections are absent in the source and flagged, not imputed).

Key columns in forecasts

  • β€”Identity: model, match_id, phase (group/knockout), stage, team_a, team_b
  • β€”Truth: outcome ∈ {teamawin, draw, teambwin} (90-minute result; a draw for the four penalty ties), decided_by, advanced
  • β€”Forecast: p_a, p_draw, p_b, bet_pick, bet_stake, reasoning
  • β€”Reflection: refl_outcome, refl_calibrated, refl_luck, refl_missed, refl_overweighted, refl_bet_diff, refl_conf
  • β€”Market & P&L: imp_home/imp_draw/imp_away, odds_*_dec, is_bet, profit, contrarian

Load it

python
from datasets import load_dataset
ds = load_dataset("dingjiacheng/wc2026-agents", "forecasts")       # default
odds = load_dataset("dingjiacheng/wc2026-agents", "odds")

FAQ

Which AI predicted the 2026 World Cup best? It depends on the metric: ChatGPT and Grok tied on accuracy (68.3%), Claude (Brier 0.4705) and Grok (0.4706) had the best probability scores, and Grok made the most virtual betting profit (+$650). None beat the betting market's Brier score (0.469).

Did any AI beat the bookmakers? No, not convincingly. The market's Brier score was better than every agent's, and backing the market favourite every match (+$1,041) out-earned all four.

Which matches did every AI get wrong? 32 of 104, 22 of them draws. In the knockout stage: Germany 1-1 Paraguay, Netherlands 1-1 Morocco, Australia 1-1 Egypt, Switzerland 0-0 Colombia (all on penalties), Brazil 1-2 Norway, USA 1-4 Belgium, France 0-2 Spain, France 4-6 England.

Did AI predict Spain's win in the final? All four agents rated a Spain win at least as likely as any other outcome before Spain beat Argentina 1-0 (Spain win probability 35% to 44%).

Ethics

All betting is virtual; the dataset is a measurement instrument, not gambling advice. Records are model outputs and public fixtures, scores and odds; no human subjects or personal data.

Citation

bibtex
@misc{ding2026wc2026agents,
  title         = {{FIFA} World Cup 2026 as a Contamination-Free Benchmark for
                   {LLM} Forecasting Agents: Four Models, a Bookmaker, and 104 Matches},
  author        = {Ding, Jiacheng and Guo, Cong},
  year          = {2026},
  eprint        = {2607.17765},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2607.17765}
}