FINAL-Bench/all-bench-leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
<p align="center"> <a href="https://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard"><img src="https://img.shields.io/badge/🏆LiveLeaderboard-ALL_Bench-6366f1?style=for-the-badge" alt="Live Leaderboard"></a> </p>
<p align="center"> <a href="https://github.com/final-bench/ALL-Bench-Leaderboard"><img src="https://img.shields.io/badge/GitHub-Repo-black?style=flat-square&logo=github" alt="GitHub"></a> <a href="https://huggingface.co/datasets/FINAL-Bench/Metacognitive"><img src="https://img.shields.io/badge/🧬FINALBench-Dataset-blueviolet?style=flat-square" alt="FINAL Bench"></a> <a href="https://huggingface.co/spaces/FINAL-Bench/Leaderboard"><img src="https://img.shields.io/badge/🧬FINALBench-Leaderboard-teal?style=flat-square" alt="FINAL Leaderboard"></a> </p>
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and decision-makers who need a trustworthy, unified view of the AI model landscape.
What's New — v2.2.1
🏅 Union Eval ★NEW
ALL Bench's proprietary integrated benchmark. Fuses the discriminative core of 10 existing benchmarks (GPQA, AIME, HLE, MMLU-Pro, IFEval, LiveCodeBench, BFCL, ARC-AGI, SWE, FINAL Bench) into a single 1000-question pool with a season-based rotation system.
Key features:
- 100% JSON auto-graded — every question requires mandatory JSON output with verifiable fields. Zero keyword matching.
- Fuzzy JSON matching — tolerates key name variants, fraction formats, text fallback when JSON parsing fails.
- Season rotation — 70% new questions each season, 30% anchor questions for cross-season IRT calibration.
- 8 rounds of empirical testing — v2 (82.4%) → v3 (82.0%) → Final (79.5%) → S2 (81.8%) → S3 (75.0%) → Fuzzy (69.9/69.3%).
Key discovery: "The bottleneck in benchmarking is not question difficulty — it's grading methodology."
Empirically confirmed LLM weakness map:
- 🔴 Poetry + code cross-constraints: 18-28%
- 🔴 Complex JSON structure (10+ constraints): 0%
- 🔴 Pure series computation (Σk²/3ᵏ): 0%
- 🟢 Metacognitive reasoning (Bayes, proof errors): 95%
- 🟢 Revised science detection: 86%
Current scores (S3, 20Q sample, Fuzzy JSON):
Other v2.2 changes
- Fair Coverage Correction: composite scoring ^0.5 → ^0.7
- +7 FINAL Bench scores (15 total)
- Columns sorted by fill rate
- Model Card popup (click model name) · FINAL Bench detail popup (click Metacog score)
- 🔥 Heatmap, 💰 Price vs Performance scatter tools
Live Leaderboard
👉 [https://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard](https://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard)
Interactive features: composite ranking, dark mode, advanced search (GPQA > 90 open, price < 1), Model Finder, Head-to-Head comparison, Trust Map heatmap, Bar Race animation, Model Card popup, FINAL Bench detail popup, and downloadable Intelligence Report (PDF/DOCX).
Data Structure
data/
├── llm.jsonl # 41 LLMs × 32 fields (incl. unionEval ★NEW)
├── vlm_flagship.jsonl # 11 flagship VLMs × 10 benchmarks
├── agent.jsonl # 10 agent models × 8 benchmarks
├── image.jsonl # 10 image gen models × S/A/B/C ratings
├── video.jsonl # 10 video gen models × S/A/B/C ratings
└── music.jsonl # 8 music gen models × S/A/B/C ratingsLLM Field Schema
Composite Score
Score = Avg(confirmed benchmarks) × (N/10)^0.710 core benchmarks across the 5-Axis Intelligence Framework: Knowledge · Expert Reasoning · Abstract Reasoning · Metacognition · Execution.
v2.2 change: Exponent adjusted from 0.5 to 0.7 for fairer coverage weighting. Models with 7/10 benchmarks receive ×0.79 (was ×0.84), while 4/10 receives ×0.53 (was ×0.63).
Confidence System
Each benchmark score in the confidence object is tagged:
Example:
"Claude Opus 4.6": {
"gpqa": { "level": "cross-verified", "source": "Anthropic + Vellum + DataCamp" },
"arcAgi2": { "level": "cross-verified", "source": "Vellum + llm-stats + NxCode + DataCamp" },
"metacog": { "level": "single-source", "source": "FINAL Bench dataset" },
"unionEval": { "level": "single-source", "source": "Union Eval S3 — ALL Bench official" }
}Usage
from datasets import load_dataset
# Load LLM data
ds = load_dataset("FINAL-Bench/ALL-Bench-Leaderboard", "llm")
df = ds["train"].to_pandas()
# Top 5 LLMs by GPQA
ranked = df.dropna(subset=["gpqa"]).sort_values("gpqa", ascending=False)
for _, m in ranked.head(5).iterrows():
print(f"{m['name']:25s} GPQA={m['gpqa']}")
# Union Eval scores
union = df.dropna(subset=["unionEval"]).sort_values("unionEval", ascending=False)
for _, m in union.iterrows():
print(f"{m['name']:25s} Union Eval={m['unionEval']}")Union Eval — Integrated AI Assessment
Union Eval is ALL Bench's proprietary benchmark designed to address three fundamental problems with existing AI evaluations:
- Contamination — Public benchmarks leak into training data. Union Eval rotates 70% of questions each season.
- Single-axis measurement — AIME tests only math, IFEval only instruction-following. Union Eval integrates arithmetic, poetry constraints, metacognition, coding, calibration, and myth detection.
- Score inflation via keyword matching — Traditional rubric grading gives 100% to "well-written" answers even if content is wrong. Union Eval enforces mandatory JSON output with zero keyword matching.
Structure (S3 — 100 Questions from 1000 Pool):
Note: The 100-question dataset is not publicly released to prevent contamination. Only scores are published.
FINAL Bench — Metacognitive Benchmark
FINAL Bench measures AI self-correction ability. Error Recovery (ER) explains 94.8% of metacognitive performance variance. 15 frontier models evaluated.
Changelog
Citation
@misc{allbench2026,
title={ALL Bench Leaderboard 2026: Unified Multi-Modal AI Evaluation},
author={ALL Bench Team},
year={2026},
url={https://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard}
}#AIBenchmark #LLMLeaderboard #GPT5 #Claude #Gemini #ALLBench #FINALBench #Metacognition #UnionEval #VLM #AIAgent #MultiModal #HuggingFace #ARC-AGI #AIEvaluation #VIDRAFT.net
