datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.spellcheck_punctuation_benchmarkRussian Spellcheck Benchmark is a new benchmark for spelling correction in Russian language.
It includes four datasets, each of which consists of pairs of sentences in Russian language.
Each pair embodies sentence, which may contain spelling errors, and its corresponding correction.
Datasets were gathered from various sources and domains including social networks, internet blogs, github commits,
medical anamnesis, literature, news, reviews and more.LEXam
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations.
Paper | Website & Leaderboard | GitHub Repository
🔥 News
[2026/01] Our paper has been accepted to ICLR 2026!
[2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.spellcheck_benchmarkRussian Spellcheck Benchmark is a new benchmark for spelling correction in Russian language.
It includes four datasets, each of which consists of pairs of sentences in Russian language.
Each pair embodies sentence, which may contain spelling errors, and its corresponding correction.
Datasets were gathered from various sources and domains including social networks, internet blogs, github commits,
medical anamnesis, literature, news, reviews and more.Open-LLM-Benchmark
Open-LLM-Benchmark
Dataset Description
The Open-LLM-Leaderboard tracks the performance of various large language models (LLMs) on open-style questions to reflect their true capability. The dataset includes pre-generated model answers and evaluations using an LLM-based evaluator.
License: CC-BY 4.0
Dataset Structure
An example of model response files looks as follows:
{
"question": "What is the main function of photosynthetic cells within a plant?"… See the full description on the dataset page: https://huggingface.co/datasets/Open-Style/Open-LLM-Benchmark.lm1bA benchmark corpus to be used for measuring progress in statistical language modeling. This has almost one billion words in the training data.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.bonsai-jetson-benchmark-7w
Bonsai Jetson Benchmark: 7W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 7W
Backend: llama.cpp build-jetson · CUDA · -ngl 99
Sweep: prompt in {256, 512, 1024, 2048} tok x gen in {128, 256, 512} tok · 20 reqs/combo
Context: 2560 tok · Concurrency: 1
Part of smolperfleaderboard, a public on-device
LLM benchmark leaderboard. Headline metric is output tok/J (tokens per joule),
computed over the decode phase.
Files
Bonsai-*, Ternary-Bonsai-*: per-combo… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-7w.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.bonsai-jetson-benchmark-25w
Bonsai Jetson Benchmark — 25W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 25W
Status: Partial — 36 combos
Bonsai All-Model Benchmark — Jetson Orin Nano Super 8GB — 25W
Date: 2026-05-27 22:38Backend: llama.cpp (build-jetson) / CUDA / -ngl 99Platform: NVIDIA Jetson Orin Nano Super 8GB (6-core Cortex-A78AE + Ampere GPU)Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok × 20 reqs/comboKey metric: tok/J = output tokens per second ÷… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-25w.bonsai-jetson-benchmark-15w
Bonsai Jetson Benchmark — 15W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W
Backend: llama.cpp build-jetson · CUDA · -ngl 99
Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo
Status: Complete — 57 combos (5 models × 12 prompt/gen configs)
Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W)
Models
Model
Quant
Size
Bonsai-1.7B
Q1_0 (1-bit)
~237 MB
Bonsai-4B
Q1_0 (1-bit)
~540 MB
Bonsai-8B
Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.bonsai-jetson-benchmark-maxn
Bonsai Jetson Benchmark — MAXN_SUPER
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: MAXN_SUPER
Status: Complete — 57 combos
Bonsai All-Model Benchmark — Jetson Orin Nano Super 8GB — MAXN_SUPER
Date: 2026-05-27 00:02Backend: llama.cpp (build-jetson) / CUDA / -ngl 99Platform: NVIDIA Jetson Orin Nano Super 8GB (6-core Cortex-A78AE + Ampere GPU)Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok × 20 reqs/comboKey metric: tok/J = output… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-maxn.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.SWE-QA-Benchmark
SWE-QA Benchmark
A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories.
Dataset Summary
Total Questions: 720
Repositories: 15
Format: JSONL (JSON Lines)
Fields: question, answer
Repository Coverage
Each repository contains 48 questions:
astropy
conan
django
flask
matplotlib
pylint
pytest
reflex
requests
scikit-learn
sphinx
sqlfluff
streamlink
sympy
xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.helium-market-resolution-benchmark
What is this?
Market Resolution contains 299 ranked option-contract questions. Most test calculations or comparisons from frozen quotes. Sixty test an implied-volatility prior with the premium hidden, and 11 test probability-of-finishing-in-the-money forecasts against a later outcome.
It does not measure trading profitability. It measures bounded option-chain reasoning: implied volatility (IV), delta, time value, parity, term structure, relative IV, chain surfaces, and… See the full description on the dataset page: https://huggingface.co/datasets/HeliumTrades/helium-market-resolution-benchmark.benchmark-radar
Benchmark Radar Dataset
Overview
Benchmark Radar is a living registry, search engine, and discovery pipeline
for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the
Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and
Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115).
As described in the paper's Two Input Paths framework, Benchmark Radar
combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.chess_puzzle_benchmark
Chess Puzzle Benchmark
Chess puzzle evaluation sets at five difficulty tiers, B1 (easiest) through
B5 (hardest). Each example is a chess game given in PGN move notation; the
model must produce the next move(s).
Two prompt variants
The same puzzles are released in two forms that differ only in the prompt suffix:
think/ — the prompt ends with a special <T> token. <T> is a
reasoning trigger: it tells the model to think (produce a chain of reasoning)
before… See the full description on the dataset page: https://huggingface.co/datasets/pavelslab-nyu/chess_puzzle_benchmark.edgar-forecast-benchmark
EDGAR-Forecast Benchmark
EDGAR-Forecast is a closed-sandbox benchmark for filing-grounded numerical forecasting from historical SEC filings in EDGAR. The benchmark contains 50 company-level instances and 250 numeric forecast targets from hidden 2026 10-Q filings.
Questions that mention 2025 refer to values disclosed in 2026 Q1 filings; those filings were filed in Q1 2026, so they remain outside the evaluated models' knowledge cutoffs.
Each benchmark directory includes the question… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/edgar-forecast-benchmark.last-translation-benchmark
Last Translation Benchmark
Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases.
Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived).
Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.helium-model-worldview-benchmark
What is this?
Model Worldview is a 323-item probe suite. It includes 93 paired tests that hold a scenario fixed while changing a source, identity, or framing cue, plus standalone value tradeoffs, political survey items, and evidence questions.
It is not a single left-right score or a ranking of the "best worldview."
Across 16 matched stereotype-essay prompts, 6 models triggered the refusal classifier every time. The lowest rate was grok-4.20-reasoning: 3/16. Same requests… See the full description on the dataset page: https://huggingface.co/datasets/HeliumTrades/helium-model-worldview-benchmark.Agent-Failure-Recovery-Benchmark
Agent Failure Recovery Benchmark
22,573 synthetic, source-verified failure → recovery trajectories across four domains.
Evaluate whether a system can reject a failed plan, choose a recovery, or recognize that no recovery exists. Every row links to a hashed record in a pinned public source and is checked by independent computational oracles and trajectory replay.
Tasks: failure diagnosis, recovery-action prediction, planning regression, scoped agent evaluation and imitation… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Agent-Failure-Recovery-Benchmark.delulu-fim-benchmarkDelulu — Fill-in-the-Middle Code Hallucination Benchmark
A verified multilingual benchmark for code-completion hallucinations.
Every golden completion compiles. Every hallucination provably doesn't.
📄 Read the preprint on arXiv →
Every Delulu sample ships as a self-contained Docker image. The viewer above lets you browse the dataset, pull a sample's verifier, and re-run verify golden / verify hallucinated / verify patch <your-completion> with one… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delulu-fim-benchmark.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.wallet-eval-benchmark
Wallet tool-calling eval benchmark
The eval side of the wallet fine-tuning work: what the models are scored on.
The training rows are deliberately not published. These cases are held out from
them by construction, and that is the only reason a score here means anything. If you
train on this benchmark, say so — a number from a contaminated run is not comparable to
the ones below.
The model these cases were used to select is public:
ef-dai-team/gemma-4-E4B-wallet-ft-v5,
which… See the full description on the dataset page: https://huggingface.co/datasets/ef-ai/wallet-eval-benchmark.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.os-omni-benchmark
OS-Omni Benchmark
OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks.
Contents
data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation.
data/tasks.jsonl: JSON Lines copy of the same task index.
metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/OSOmni/os-omni-benchmark.eulerfold-benchmark
EulerFold Technical & Scientific Benchmark Dataset
The EulerFold Benchmark is a unified, multi-format academic dataset containing 10,017 verified technical problems mapped directly across 10 foundational domains, 329 subject areas, and 3,030 standardized subtopics.
Designed to evaluate complex reasoning and power adaptive technical learning systems, the benchmark spans multiple-choice diagnostics with structured misconception maps, open-ended mathematical derivations and… See the full description on the dataset page: https://huggingface.co/datasets/csankalp21/eulerfold-benchmark.TACO-Benchmark
TACO-Benchmark
TACO (Text-to-SQL with Ambiguous and Cross-database Open-domain queries) is a benchmark for real-world data-lake Text-to-SQL.
📢 News (2026): TACO has been accepted to VLDB 2026! 🎉📄 Paper: arXiv:2606.14201
GitHub (code & evaluation): Akanezora0/TACO-Benchmark
Google Drive mirror: TACO-Benchmark.zip
Overview
Unlike Spider or BIRD — where the target database is known and schemas are clean — TACO evaluates systems on three challenges common in… See the full description on the dataset page: https://huggingface.co/datasets/Akanezora/TACO-Benchmark.
