Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmarena-ai /leaderboard-dataset Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load_dataset # Load all historical text style control data ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full") # Load the current text style control leaderboard ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest") # Filter to overall category ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.tabular1M<n<10M28 likes72k downloads3d agoHugging Face02qimma /leaderboard-detailstext1M<n<10M0 likes32k downloads1mo agoHugging Face03cot-leaderboard /cot-eval-traces-2.0text1M<n<10M9 likes14k downloads2y agoHugging Face04open-llm-leaderboard /contentstabular1K<n<10K25 likes14k downloads2y agoHugging Face05nebius /SWE-rebench-leaderboard Dataset Summary ❗❗❗ Please use Harbour Hub for the July 2026 evaluation split:https://hub.harborframework.com/datasets/ibragim-badertdinov/swe-rebench-07-2026/latest SWE-rebench-leaderboard is a continuously updated, curated subset of the full SWE-rebench corpus, tailored for benchmarking software engineering agents on real-world tasks. These tasks are used in the SWE-rebench leaderboard. For more details on the benchmark methodology and data collection process, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-leaderboard.tabular1K<n<10K30 likes7.4k downloads2mo agoHugging Face06nithinraok /asr-leaderboard-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.audioautomatic-speech-recognition100K<n<1M4 likes4.4k downloads1y agoHugging Face07RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M5 likes2.8k downloads4h agoHugging Face08Mohaaxa /quantbench-leaderboard-data QuantBench leaderboard data Raw benchmark data behind the QuantBench leaderboard: calibration-quality GPTQ/AWQ quantization results across model sizes, calibration corpora, and GPU tiers. 331 rows (239 ok / 92 failed — failed runs are published too; a documented failure is a finding, not noise). Models Qwen/Qwen2.5-1.5B-Instruct (1.5B) HuggingFaceTB/SmolLM2-1.7B-Instruct (1.7B) deepgrove/Bonsai (0.5B) Qwen/Qwen2.5-3B-Instruct (3B) — licence pending, rows only, no… See the full description on the dataset page: https://huggingface.co/datasets/Mohaaxa/quantbench-leaderboard-data.tabularn<1K0 likes2.3k downloads2mo agoHugging Face09OpenEvals /leaderboard-datatabularn<1K1 likes2.1k downloads6mo agoHugging Face10huggingworld /open-asr-leaderboard ESB Test Sets: Parquet & Sorted This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length. The format is also changed, from custom loading script (un-safe remote code) to parquet (safe). Broadly speaking, this dataset was generated with the following code-snippet: from datasets import load_dataset, get_dataset_config_names DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/huggingworld/open-asr-leaderboard.audio100K<n<1M1 likes1.7k downloads5mo agoHugging Face11open-rl-leaderboard /results_v2 Dataset Card for "results_v2" Leaderboard More Information needed text10M<n<100M1 likes1.6k downloads2y agoHugging Face12ioi-leaderboard /ioi-eval-dummy-openrouter_openai_gpt-3.5-turbotextn<1K0 likes1.6k downloads2y agoHugging Face13ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbotextn<1K0 likes1.6k downloads2y agoHugging Face14ioi-leaderboard /ioi-eval-openrouter_openai_o1textn<1K0 likes1.6k downloads2y agoHugging Face15ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbo-new-prompttextn<1K0 likes1.6k downloads2y agoHugging Face16ioi-leaderboard /ioi-eval-openrouter_anthropic_claude-3_7-sonnet_thinking-prompt-mem-limittextn<1K0 likes1.6k downloads2y agoHugging Face17ioi-leaderboard /ioi-eval-openrouter_openai_o3-minitextn<1K0 likes1.6k downloads2y agoHugging Face18ioi-leaderboard /ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limittextn<1K0 likes1.6k downloads2y agoHugging Face19ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbo-prompt-mem-limittextn<1K0 likes1.6k downloads2y agoHugging Face20ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbo-texttextn<1K0 likes1.6k downloads2y agoHugging Face21ioi-leaderboard /ioi-eval-openrouter_openai_o1-minitextn<1K0 likes1.5k downloads2y agoHugging Face22llm-jp /leaderboard-contents-v2tabularn<1K1 likes1.5k downloads15d agoHugging Face23hf-audio /open-asr-leaderboard-multilingual-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition100K<n<1M4 likes1.4k downloads3mo agoHugging Face24GSMA /leaderboard Open Telco Leaderboard Scores Benchmark scores for 84 models across 7 telecom-domain benchmarks, sourced from the MWC leaderboard. This dataset publishes scores only (no energy metrics). Files leaderboard_scores.csv: Flat table for the dataset viewer. leaderboard_scores.json: Structured JSON with per-model benchmark scores and standard errors. Schema (leaderboard_scores.csv) Core columns: model — Model name provider — Model provider (e.g. OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/leaderboard.tabulartext-classificationn<1K6 likes1.3k downloads16h agoHugging Face25VoiceArena /MonsoonASR-Open-ASR-leaderboard-en-IN Voice Arena Monsoon en-IN (public test) Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed. A conversational Indian English ASR test set that records who was speaking, not only what was said. Every clip carries twelve speaker attributes — gender, age, native district and state, education, occupation, income band, handset — so a difference between two systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.audioautomatic-speech-recognition1K<n<10K4 likes967 downloads1mo agoHugging Face26quintelamanuel /political-leaderboard-resultstabularn<1K2 likes495 downloads4mo agoHugging Face27MyHeartCounts /OpenMHC-leaderboard-data OpenMHC Leaderboard Data Per-user substrate behind the OpenMHC wearable-health benchmark leaderboard. Each file is one method's reduced per-user, per-task values for one track; the leaderboard recompute consumes these to produce paired skill scores, cross-method ranks, and fairness skill scores. This repo holds reduced metrics / predictions keyed by pseudonymous participant id — not raw sensor data. Layout <track>/<method>.parquet e.g.… See the full description on the dataset page: https://huggingface.co/datasets/MyHeartCounts/OpenMHC-leaderboard-data.tabular10M<n<100M1 likes345 downloads2mo agoHugging Face28simonycl /persuasiveness-leaderboard-invertedtext10K<n<100K0 likes339 downloads11mo agoHugging Face29SaarAI /asr-leaderboard-datasetsgated Afrivoice and Amharic configs These configs were added by _data-prep-gsma/prep_final.py and _data-prep-gsma/prep_amharic.py in the gsma-asr-bench project. All Afrivoice rows are test-only: each config below contains exactly the held-out evaluation partition from the upstream dataset. Common schema (identical across the three configs): column type notes file_name string stable per-config identifier audio Audio(sampling_rate=16000) mono duration float64 seconds… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-leaderboard-datasets.audio100K<n<1M0 likes296 downloads7d agoHugging Face30stacklok /llm-security-leaderboard-contentstabularn<1K0 likes294 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.