datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.regx-benchmark
RegX
Cross-Domain Multi-View Point Cloud Registration Benchmark
RegX evaluates multi-view point cloud registration across scales spanning nine orders
of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors
never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and
solid-state LiDAR, terrestrial and airborne laser scanners.
Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question
instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.ai-humanizer-benchmark
AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026)
AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.invoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.hack-ignition-benchmark
hack-ignition benchmark — data, v0.1.6
Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL
comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt,
training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores
what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.benchmark-radar
Benchmark Radar Dataset
Overview
Benchmark Radar is a living registry, search engine, and discovery pipeline
for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the
Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and
Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115).
As described in the paper's Two Input Paths framework, Benchmark Radar
combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.scene-mem-benchmark
scene-mem-benchmark
A benchmark for scene memory in embodied agents: an agent watches a mobile manipulator work
in a house for several minutes, then is asked to retrieve an object it has to remember — one
that was moved, dropped, or merely seen along the way — or (resume) to go back and finish the
job it was interrupted in, remembering how far it had got — or (routine) to put a new object away
where this household keeps that kind of thing, a rule it was never told and can only… See the full description on the dataset page: https://huggingface.co/datasets/Keh0t0/scene-mem-benchmark.data-product-benchmark
DPDisc Dataset
Paper | Code
Dataset Description
This dataset provides a benchmark for automatic data product creation. The task is framed as follows: given a natural language data product request and a corpus of text and tables, the objective is to identify the relevant tables and text documents that should be included in the resulting data product which would useful to the given data product request. The benchmark brings together three variants: HybridQA, TAT-QA, and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/data-product-benchmark.mea-benchmark
MEA-Benchmark
Benchmark for MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations.
📄 Paper: arXiv:2610.02480
💻 Code: github.com/AikyamLab/xai-agent
MEA-Benchmark evaluates explanations of neural network models across three modalities (tabular, text, vision) with ten question types (Q1–Q10), spanning feature attribution, counterfactual reasoning, and spurious feature detection. Each question type is paired with a perturbation-based faithfulness metric (see… See the full description on the dataset page: https://huggingface.co/datasets/EstherrrCheng/mea-benchmark.Finers-4k-benchmarklegal-llm-benchmark
Legal LLM Benchmark Dataset
Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models
Quick Start
from datasets import load_dataset
# Core datasets
questions = load_dataset("marvintong/legal-llm-benchmark", "questions")
phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations")
phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations")
# Additional helpful datasets
contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.rag_instruct_benchmark_tester
Dataset Card for RAG-Instruct-Benchmark-Tester
Dataset Summary
This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases,
contracts, invoices, technical articles, general news and short texts.
The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.TrialPanorama-benchmarkDataset website: https://ryanwangzf.github.io/projects/trialpanorama
knows-benchmark
KNOWS Benchmark
KNOWS evaluates web agents on the work people actually do in Google Workspace: writing documents,
building spreadsheets, and composing slide decks that require web research, multi-step tool use, and
faithful grounding in retrieved sources.
This dataset contains the task definitions — the prompt an agent receives, plus the structured
evaluation rubric used to grade the artifact it produces.
Tasks
110 (22 templates × 5 instances)
Domains
20
Mean… See the full description on the dataset page: https://huggingface.co/datasets/utahnlp/knows-benchmark.ai-benchmarks-v2-2026
Ai Benchmarks V2 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-benchmarks-v2-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-benchmarks-v2-2026.SciCode-Runnable-Benchmark-Reviewedstreaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.UltraMS-Benchmark-Assets
UltraMS benchmark assets
Inputs for reproducing the UltraMS spectral-property and molecular-identification benchmarks. The GitHub benchmark guides provide the training and evaluation commands.
Path
Use
spectral_properties/masked_peak_reconstruction.pt
UltraMS checkpoint for masked peak reconstruction.
spectral_properties/massspecgym_labels.parquet
Frozen MassSpecGym property and neutral-loss training table used in Figure 2.
spectral_properties/heteroatom_count/… See the full description on the dataset page: https://huggingface.co/datasets/dsadd4/UltraMS-Benchmark-Assets.multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.vqa-cmsv-benchmark
VQA-CMSV Benchmark Data Package
This repository contains annotation splits for VQA v2-CMSV, GQA-CMSV, and VG-CMSV, plus patch-mask NPZ files used for mask supervision experiments.
Contents
data/vqa_v2_cmsv/train.json, data/vqa_v2_cmsv/val.json, data/vqa_v2_cmsv/test.json
data/gqa_cmsv/train.jsonl, data/gqa_cmsv/val.jsonl, data/gqa_cmsv/test.jsonl
data/vg_cmsv/train.jsonl, data/vg_cmsv/val.jsonl, data/vg_cmsv/test.jsonl
masks/vqa_v2_cmsv_masks.npz
masks/gqa_cmsv_masks.npz… See the full description on the dataset page: https://huggingface.co/datasets/as-benchmark-artifacts/vqa-cmsv-benchmark.SKILLRET
SkillRet Benchmark
SkillRet is a retrieval benchmark for matching natural-language user requests to
agent skills. Each retrieval document is a full agent skill, represented by its
name, short description, and full Markdown skill body. Each query describes a
realistic user request that requires one or more relevant skills.
The benchmark is built from public agent skills indexed from GitHub and contains
synthetic train and evaluation queries generated through a self-instruct-style… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-ed-benchmark/SKILLRET.MEME
MEME: Multi-Entity and Evolving Memory Evaluation
A benchmark for evaluating LLM memory systems along two orthogonal dimensions: entity scope (single vs. multi-entity) and temporal dynamics (static vs. evolving). MEME defines six tasks targeting memory-intensive operations in each quadrant, including two task types that no prior benchmark covers: Cascade (propagating updates through dependency rules) and Absence (recognizing uncertainty when a previously valid answer becomes… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME.Benchmark-Testingcompile-benchmark
CompilingThings Compile Benchmark for MQL5®
This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.mcphunt-agent-traces
MCPHunt Agent Traces
Agent execution traces from the MCPHunt evaluation framework, measuring
cross-boundary data propagation in multi-server MCP agents.
Contents
main/ — 3,615 traces from 5 models across 147 tasks and 7 environment
variants (risky_v1/v2/v3, benign, hard_neg_v1/v2/v3). One JSON file per model.
mitigation/ — 2,706 traces from the prompt-mitigation study (M0--M3
levels) across 3 models.
live_guard_defense/ — 387 DeepSeek-V4-Flash traces from the… See the full description on the dataset page: https://huggingface.co/datasets/mcphunt-benchmark/mcphunt-agent-traces.hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
Altworld developed and published
Hemmingway-1. Bobby Pierce
published these quantizations and the evaluation package. The
collection
links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching
across reversed packets. Read CORRECTION.md before using the
aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces
Agent traces
Agent sessions published from a Trackio Logbook.
powergrid-benchmark2
Note: This is a dummy dataset for browsing only. The official dataset can be downloaded from our pipeline (Data Hub).
peptide-reasoning-benchmark
Peptide Reasoning Benchmark
PEB v1.0-RC benchmark release for peptide-reasoning model evaluation.
Includes cases, splits, baselines, references, and leaderboard artifacts.
GitHub: https://github.com/ray-r-ren/peptide-reasoning-bench
Trained a small reference LoRA model: https://huggingface.co/rayrren/the-spice-v0-mvp
