datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard-dataset
Arena Leaderboard Dataset
Historical snapshots of the Arena leaderboard.
Usage
from datasets import load_dataset
# Load all historical text style control data
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full")
# Load the current text style control leaderboard
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest")
# Filter to overall category
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.leaderboard-detailscot-eval-traces-2.0contentsSWE-rebench-leaderboard
Dataset Summary
❗❗❗ Please use Harbour Hub for the July 2026 evaluation split:https://hub.harborframework.com/datasets/ibragim-badertdinov/swe-rebench-07-2026/latest
SWE-rebench-leaderboard is a continuously updated, curated subset of the full SWE-rebench corpus, tailored for benchmarking software engineering agents on real-world tasks.
These tasks are used in the SWE-rebench leaderboard. For more details on the benchmark methodology and data collection process, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-leaderboard.asr-leaderboard-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.quantbench-leaderboard-data
QuantBench leaderboard data
Raw benchmark data behind the QuantBench leaderboard:
calibration-quality GPTQ/AWQ quantization results across model sizes, calibration
corpora, and GPU tiers. 331 rows (239 ok / 92 failed — failed
runs are published too; a documented failure is a finding, not noise).
Models
Qwen/Qwen2.5-1.5B-Instruct (1.5B)
HuggingFaceTB/SmolLM2-1.7B-Instruct (1.7B)
deepgrove/Bonsai (0.5B)
Qwen/Qwen2.5-3B-Instruct (3B) — licence pending, rows only, no… See the full description on the dataset page: https://huggingface.co/datasets/Mohaaxa/quantbench-leaderboard-data.leaderboard-dataopen-asr-leaderboard
ESB Test Sets: Parquet & Sorted
This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length.
The format is also changed, from custom loading script (un-safe remote code) to parquet (safe).
Broadly speaking, this dataset was generated with the following code-snippet:
from datasets import load_dataset, get_dataset_config_names
DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from
HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/huggingworld/open-asr-leaderboard.results_v2
Dataset Card for "results_v2"
Leaderboard
More Information needed
ioi-eval-dummy-openrouter_openai_gpt-3.5-turboioi-eval-openrouter_openai_gpt-3.5-turboioi-eval-openrouter_openai_o1ioi-eval-openrouter_openai_gpt-3.5-turbo-new-promptioi-eval-openrouter_anthropic_claude-3_7-sonnet_thinking-prompt-mem-limitioi-eval-openrouter_openai_o3-miniioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limitioi-eval-openrouter_openai_gpt-3.5-turbo-prompt-mem-limitioi-eval-openrouter_openai_gpt-3.5-turbo-textioi-eval-openrouter_openai_o1-minileaderboard-contents-v2open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.leaderboard
Open Telco Leaderboard Scores
Benchmark scores for 84 models across 7 telecom-domain benchmarks, sourced from the MWC leaderboard.
This dataset publishes scores only (no energy metrics).
Files
leaderboard_scores.csv: Flat table for the dataset viewer.
leaderboard_scores.json: Structured JSON with per-model benchmark scores and standard errors.
Schema (leaderboard_scores.csv)
Core columns:
model — Model name
provider — Model provider (e.g. OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/leaderboard.MonsoonASR-Open-ASR-leaderboard-en-IN
Voice Arena Monsoon en-IN (public test)
Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed.
A conversational Indian English ASR test set that records who was speaking, not only what
was said. Every clip carries twelve speaker attributes — gender, age, native district
and state, education, occupation, income band, handset — so a difference between two
systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.political-leaderboard-resultsOpenMHC-leaderboard-data
OpenMHC Leaderboard Data
Per-user substrate behind the OpenMHC wearable-health benchmark leaderboard. Each file is one method's reduced per-user, per-task values for one track; the leaderboard recompute consumes these to produce paired skill scores, cross-method ranks, and fairness skill scores.
This repo holds reduced metrics / predictions keyed by pseudonymous participant id — not raw sensor data.
Layout
<track>/<method>.parquet e.g.… See the full description on the dataset page: https://huggingface.co/datasets/MyHeartCounts/OpenMHC-leaderboard-data.persuasiveness-leaderboard-invertedasr-leaderboard-datasets
Afrivoice and Amharic configs
These configs were added by _data-prep-gsma/prep_final.py and
_data-prep-gsma/prep_amharic.py in the gsma-asr-bench
project. All Afrivoice rows are test-only: each config below contains
exactly the held-out evaluation partition from the upstream dataset.
Common schema (identical across the three configs):
column
type
notes
file_name
string
stable per-config identifier
audio
Audio(sampling_rate=16000)
mono
duration
float64
seconds… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-leaderboard-datasets.llm-security-leaderboard-contents
