datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard-dataset
Arena Leaderboard Dataset
Historical snapshots of the Arena leaderboard.
Usage
from datasets import load_dataset
# Load all historical text style control data
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full")
# Load the current text style control leaderboard
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest")
# Filter to overall category
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.drlc-leaderboard-datacontentsopen-asr-leaderboard-resultsSWE-rebench-leaderboard
Dataset Summary
❗❗❗ Please use Harbour Hub for the July 2026 evaluation split:https://hub.harborframework.com/datasets/ibragim-badertdinov/swe-rebench-07-2026/latest
SWE-rebench-leaderboard is a continuously updated, curated subset of the full SWE-rebench corpus, tailored for benchmarking software engineering agents on real-world tasks.
These tasks are used in the SWE-rebench leaderboard. For more details on the benchmark methodology and data collection process, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-leaderboard.LHTB-leaderboard
LHTB Leaderboard — Long-Horizon Terminal-Bench
This repository hosts submitted runs for
Long-Horizon Terminal-Bench (LHTB),
a 46-task benchmark measuring how well LLM agents sustain useful work in a
containerized terminal over hundreds of steps.
Every entry below ships its complete run artifacts — per-trial configs, results,
verifier outputs and terminal recordings — so any score on this board can be audited
without rerunning the suite.
📊 Benchmark dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.quantbench-leaderboard-data
QuantBench leaderboard data
Raw benchmark data behind the QuantBench leaderboard:
calibration-quality GPTQ/AWQ quantization results across model sizes, calibration
corpora, and GPU tiers. 331 rows (239 ok / 92 failed — failed
runs are published too; a documented failure is a finding, not noise).
Models
Qwen/Qwen2.5-1.5B-Instruct (1.5B)
HuggingFaceTB/SmolLM2-1.7B-Instruct (1.7B)
deepgrove/Bonsai (0.5B)
Qwen/Qwen2.5-3B-Instruct (3B) — licence pending, rows only, no… See the full description on the dataset page: https://huggingface.co/datasets/Mohaaxa/quantbench-leaderboard-data.leaderboard-dataleaderboard_longformquantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.leaderboard-contents-v2leaderboard
Open Telco Leaderboard Scores
Benchmark scores for 84 models across 7 telecom-domain benchmarks, sourced from the MWC leaderboard.
This dataset publishes scores only (no energy metrics).
Files
leaderboard_scores.csv: Flat table for the dataset viewer.
leaderboard_scores.json: Structured JSON with per-model benchmark scores and standard errors.
Schema (leaderboard_scores.csv)
Core columns:
model — Model name
provider — Model provider (e.g. OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/leaderboard.chat-resultsALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.armnet-demo-leaderboardllm-hkmmlu-leaderboard-requestsbergson-wikitext-gpt2-leaderboard-bank
bergson leaderboard: retrain banks, scores and LDS/QLD results (WikiText GPT-2)
Everything behind the numbers on the bergson leaderboard,
for the model at EleutherAI/bergson-wikitext-gpt2-leaderboard.
path
what it is
bank/
the LDS ground truth: 100 random leave-1%-out subsets of the 4,608 training chunks (subsets.json) and each subset's measured loss change on the 50 test queries (validation.csv)
random/retrained/{base,subset_0..99}
the retrained models themselves… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-gpt2-leaderboard-bank.drlc-leaderboard-datavntl-leaderboard
VNTL Leaderboard
The VNTL leaderboard ranks Large Language Models (LLMs) based on their performance in translating Japanese Visual Novels into English. Please be aware that the current results are preliminary and subject to change as new models are evaluated, or changes are done in the evaluation script.
Comparison with Established Translation Tools
For comparison, this table shows the scores for established translation tools. These include both widely available online… See the full description on the dataset page: https://huggingface.co/datasets/lmg-anon/vntl-leaderboard.requestsleaderboard-requests
mauroibz/leaderboard-requests
Evaluation requests for the leaderboard
This dataset contains evaluation requests submitted to the leaderboard system.
Structure
Each JSON file represents an evaluation request
Files contain model information and submission metadata
Status field indicates the current state of the evaluation
Usage
These requests are used by the leaderboard system to track evaluation submissions.
ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.political-leaderboard-resultsevaluator-leaderboardOpenMHC-leaderboard-data
OpenMHC Leaderboard Data
Per-user substrate behind the OpenMHC wearable-health benchmark leaderboard. Each file is one method's reduced per-user, per-task values for one track; the leaderboard recompute consumes these to produce paired skill scores, cross-method ranks, and fairness skill scores.
This repo holds reduced metrics / predictions keyed by pseudonymous participant id — not raw sensor data.
Layout
<track>/<method>.parquet e.g.… See the full description on the dataset page: https://huggingface.co/datasets/MyHeartCounts/OpenMHC-leaderboard-data.llm-leaderboard
LLM Leaderboard (Arena x API pricing)
Daily-updated LLM leaderboard combining LMArena human-preference (Arena) scores
with OpenRouter public pricing, so you can answer "what should I use at $X?" —
a question neither source answers on its own.
Source of truth: 17nas.com/llm-leaderboard.php
| Upstream repo: AmigaMeow/llm-leaderboard-data
Files
Path
内容
data/models.jsonl
推荐入口:一行一个模型(扁平结构,可直接 load_dataset)
data/latest.json
最近一次快照的原始结构(含 sources 元信息)… See the full description on the dataset page: https://huggingface.co/datasets/AmigaMeow/llm-leaderboard.llm-security-leaderboard-contentswebgpu-bench-leaderboardsmoltrace-leaderboard
Tiny Agents. Total Visibility.
SMOLTRACE Leaderboard
This dataset contains aggregated evaluation metrics for comparing model performance across SMOLTRACE benchmark runs.
Dataset Information
Field
Value
Owner
kshitijthakkar
Updated
2026-09-29 09:28:40 UTC
Purpose
Model comparison and ranking
Schema
Identification
Column
Type
Description
run_id
string
Unique run identifier… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/smoltrace-leaderboard.
