datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.Open-LLM-Benchmark
Open-LLM-Benchmark
Dataset Description
The Open-LLM-Leaderboard tracks the performance of various large language models (LLMs) on open-style questions to reflect their true capability. The dataset includes pre-generated model answers and evaluations using an LLM-based evaluator.
License: CC-BY 4.0
Dataset Structure
An example of model response files looks as follows:
{
"question": "What is the main function of photosynthetic cells within a plant?"… See the full description on the dataset page: https://huggingface.co/datasets/Open-Style/Open-LLM-Benchmark.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.SWE-QA-Benchmark
SWE-QA Benchmark
A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories.
Dataset Summary
Total Questions: 720
Repositories: 15
Format: JSONL (JSON Lines)
Fields: question, answer
Repository Coverage
Each repository contains 48 questions:
astropy
conan
django
flask
matplotlib
pylint
pytest
reflex
requests
scikit-learn
sphinx
sqlfluff
streamlink
sympy
xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.last-translation-benchmark
Last Translation Benchmark
Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases.
Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived).
Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.benchmark-radar
Benchmark Radar Dataset
Overview
Benchmark Radar is a living registry, search engine, and discovery pipeline
for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the
Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and
Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115).
As described in the paper's Two Input Paths framework, Benchmark Radar
combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.joyo-kanji-yomi-benchmark-parakeet
日本語 | English
常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet)
常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。
このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。
このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。
概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.legal-llm-benchmark
Legal LLM Benchmark Dataset
Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models
Quick Start
from datasets import load_dataset
# Core datasets
questions = load_dataset("marvintong/legal-llm-benchmark", "questions")
phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations")
phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations")
# Additional helpful datasets
contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.mcp-agent-trajectory-benchmark
🛑 Stop LLM Agent Collapse Caused by State Corruption
Most agent failures are not tool failures. They are state failures.
The model believes the world is still valid — when reality has already changed.
This MCP trajectory dataset trains belief revision, state recovery, and autonomous replanning:
believed : 45/50 rooms synced ← agent trusts a stale worker log
tool : "success", count: 0 ← silent no-op (the worst failure mode)
verify : GDS reports actual = 40 ←… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.llmsql-benchmark
LLMSQL Benchmark
⚠️ A newer version of this dataset is available:👉 https://huggingface.co/datasets/llmsql-bench/llmsql-2.0
This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see https://github.com/LLMSQL/llmsql-benchmark.
Arxiv Article: https://arxiv.org/abs/2510.02350
Files
tables.jsonl — Database table metadata
questions.jsonl — All available questions
train_questions.jsonl, val_questions.jsonl, test_questions.jsonl — Data… See the full description on the dataset page: https://huggingface.co/datasets/llmsql-bench/llmsql-benchmark.html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/big-finance-benchmark.marketing-benchmark-of-more-than-10-ai-models
Marketing Benchmark of 10+ AI Models
A 5,000-question benchmark for evaluating LLMs across six dimensions of modern
marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical
Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas.
Every question is independently authored by the AdsGPT Marketing Bench team.
Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended
and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.SciCode-Runnable-Benchmark-Reviewedbig-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.compile-benchmark
CompilingThings Compile Benchmark for MQL5®
This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.enterprise-text-to-sql-benchmark
Enterprise Text-to-SQL Benchmark
3,087 natural-language questions paired with executable PostgreSQL, over a
12-table enterprise schema (sales, catalogue, logistics, HR).
Built to answer one question honestly: does fine-tuning actually improve
text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict
execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a
self-correction loop — and the benchmark is designed so that number cannot be
inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.LogiTraj-Benchmark
LogiTraj Benchmark
Dual license. Dataset material is licensed under CC BY 4.0;
software in evaluation/code/ is licensed under Apache-2.0. See LICENSE,
LICENSE-DATA, and LICENSE-CODE. Commercial-model raw outputs are not
included.
Synthetic Chinese enterprise logistics sandboxes, tasks, documents, versioned evaluators, verdicts, and Core/Silver/Audit quality views.
Task-family coverage is source-faithful rather than imputed: the 20260628_v45, 20260628_v46, and 20260628_v50 SFT… See the full description on the dataset page: https://huggingface.co/datasets/Linmumu009/LogiTraj-Benchmark.multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.sud-resh-benchmark
Представляем вашему вниманию бенчмарк для оценки ответов больших языковых моделей в домене российского права.
Бенчмарк создан на вычислительный грант от Yandex OpenSource
Бенчмарк основан на анонимизированных решениях судов в следующих отраслях:
Административное право
Конституционное право
Экологическое право
Финансовое право
Гражданское право
Семейное право
Право социального обеспечения
Трудовое право
Уголовное право
Жилищное право… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud-resh-benchmark.RBAC-Text2SQL-Benchmark
RBAC-Text2SQL Benchmark
Role-conditioned Text-to-SQL instances for evaluating whether LLMs generate SQL that
respects Role-Based Access Control (RBAC) constraints. Each instance pairs a natural
language question with a role policy; the model must either produce a correct SQL query
that touches only authorized resources, or refuse with Sorry, I cannot answer.
Code, evaluation harness, and reproduction instructions:
https://github.com/2020dfff/RBAC-Text2SQL-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/sharkiefff/RBAC-Text2SQL-Benchmark.UltraRAG_Benchmark
UltraRAG 2.0: Accelerating RAG for Scientific Research
UltraRAG 2.0 (UR-2.0) is jointly released by THUNLP, NEUIR, OpenBMB, and AI9Stars. It is the first lightweight RAG system construction framework built on the Model Context Protocol (MCP) architecture, designed to provide efficient modeling support for scientific research and exploration. The framework offers a full suite of teaching examples from beginner to advanced levels, integrates 17 mainstream benchmark tasks and a wide… See the full description on the dataset page: https://huggingface.co/datasets/UltraRAG/UltraRAG_Benchmark.kids-multilingual-benchmark
TinyAya v2 — Multilingual Benchmark for Children's AI Companions
2,312 child–AI conversational prompts across 23 languages, evaluated against
four models with five-judge LLM-as-judge validation.
📄 Companion article: see HF Articles by @batuhanaktas.
💻 Code: https://github.com/aktasbatuhan/cohere-tiny-aya-for-kids
Dataset summary
This dataset contains:
benchmark/items.jsonl — 2,312 benchmark items in 23 languages. Each item
is a structured prompt designed to mimic… See the full description on the dataset page: https://huggingface.co/datasets/batuhanaktas/kids-multilingual-benchmark.hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
Altworld developed and published
Hemmingway-1. Bobby Pierce
published these quantizations and the evaluation package. The
collection
links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching
across reversed packets. Read CORRECTION.md before using the
aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/big-finance-benchmark.agentic-task-benchmark
YouMind Agentic Task Benchmark
v0.1.1 · Experimental
YouMindInc/agentic-task-benchmark is a small experimental benchmark for
evaluating creative and research task outcomes against selected references.
This release contains four task descriptions, a standard outcome format, and
an offline text scorer. Task definitions and evaluation protocols may change.
Tasks
Configuration
Task
Status
image_generation
Riso portrait
Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.
