datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docmath-eval-failures-200
DocMath-Eval Failures 200: Agent Benchmark & Leaderboard
A curated benchmark of 200 challenging financial math questions that leading AI models
failed to answer correctly, with comprehensive evaluation results from multiple AI agents.
Leaderboard
Evaluated on 2026-02-21 using LLM-as-Judge (Qwen QwQ-32B) for soft scoring.
Rank
Agent
Model
Exact Match
Judge: Exact
Judge: Approx
Judge: Total
Wrong
Avg Duration
Avg Tool Calls
1
TRAE Agent
Opus 4.5
98/200 (49.0%)
96… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/docmath-eval-failures-200.aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.compute-eval
Dataset Card for ComputeEval
ComputeEval is a benchmark for evaluating LLM-generated CUDA code on correctness and performance. Each problem provides a self-contained programming challenge — spanning kernels, runtime APIs, memory management, parallel algorithms, GPU libraries, and template/DSL kernel authoring — with a held-out test harness for functional validation and optional performance benchmarks for measuring GPU execution time against a baseline.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/compute-eval.evalarc-independent-swe
Independent-source SWE workflow records
36 Qwen3-8B attempts compare four fixed workflows on three public SWE-bench
Verified tasks. There are no accepted attempts: 31 have assessable native
reports and five retain an upstream infrastructure flag, so their task outcome
is uncertain. Eight attempts produced nonempty patches. One generation request
has incomplete usage.
Inspect the interactive report
· English method · 中文方法
· Offline review and exact raw records
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/evalarc-independent-swe.llm-evaluation-self-audit
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.peerreview-bench
PeerReview Bench
CMU Paper Reviewer:https://prometheus-eval.github.io/cmu-paper-reviewer/
Repository:https://github.com/prometheus-eval/cmu-paper-reviewer
Paper:https://arxiv.org/abs/2605.20668
Point of Contact:seungone@kaist.ac.kr
Expert-annotated review items from scientific papers, organized for three
complementary evaluation tasks. All data in this dataset is intended
for evaluation, not training. All configs reference a shared, deduplicated
file store (submitted_papers)… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/peerreview-bench.long-horizon-eval
long-horizon-eval
Evaluation results for long-horizon agent performance
Dataset Description
This dataset contains evaluation results for agent trajectories, including quality assessments and performance metrics.
Dataset Structure
The dataset is organized by model name, with each model having separate JSONL files for different experimental passes.
long-horizon-eval/
├── model-1/
│ ├── pass@1.jsonl
│ ├── pass@2.jsonl
│ └── pass@3.jsonl
├── model-2/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/ToolGym/long-horizon-eval.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.Rainbow-Pony-100m-Flutter-steps-eval
Rainbow-Pony-100M Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In steps mode, the model is given an existing file and an edit instruction and
generates a sequence of localized search/replace edit actions, each mechanically
applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.rcga-evaluation-data
RCGA / LoopSFT evaluation input snapshots
Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup.
Collections
data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.r8-eval-suite-5bucket
⚠️ CRITICAL: Ollama Inference Flag Required for derived models
If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama,
you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use.
The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag.
See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned.
R8/R9 Five-Bucket… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-eval-suite-5bucket.LabCraft-Eval
LabCraft-Eval
LabCraft-Eval is an Inspect AI evaluation environment for measuring how well AI
agents execute benign molecular-microbiology protocols inside a seeded
laboratory simulator with task-dependent stochasticity. It pairs task prompts
and tool-accessible lab operations with deterministic, multi-axis trajectory
scoring.
This Hugging Face dataset export is generated from the GitHub repository:
https://github.com/jang1563/LabCraft-Eval.git
Release
Release… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/LabCraft-Eval.Rainbow-Pony-100m-Flutter-direct-eval
Rainbow-Pony-100M Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In direct mode, the model is given an existing file and an edit instruction and
generates the complete modified file in a single forward pass (as opposed to the
steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.stateless-mcp-agent-evaluation-suite-2026
⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite
A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1
⚡ Overview & Industry Problem
As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.mume-eval-suites
Mume evaluation suites
The frozen evaluation sets behind every bits-per-byte number of the Muse Mesh English and Math models (mume-english-125m, mume-math-125m), record for record, including the contamination-clean variants; and training_manifests, the list of source rows each model was trained on (no text), so the training data can be rebuilt from the public sources.
From Muse Mesh (Hugging Face). Part of the Muse Mesh English and Math models (collections). Every training run… See the full description on the dataset page: https://huggingface.co/datasets/MuseMesh/mume-eval-suites.econ-eval
econ-eval: how much do you give up by using a cheap model for an economist's work?
A reproducible benchmark of frontier and cheap LLMs on the work a trade and
policy economist actually does: Balassa RCA from raw BACI values, CAGR and
share arithmetic, bank capital and systemic-risk formulas, small trade-data
pipeline functions, checking a colleague's numbers, and policy writing in
English and Bangla. Every task carries a source field, every reference value
is derived from… See the full description on the dataset page: https://huggingface.co/datasets/deluair/econ-eval.Qwen2.5-Coder-0.5B-Flutter-direct-eval
Qwen2.5-Coder-0.5B Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In direct
mode, the model is given an existing file and an edit instruction and generates the
complete modified file in a single forward pass (as opposed to the steps /
iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct-eval.econ_eval
The Price of Progress: Benchmark-Level LLM Inference Cost Dataset
Dataset Summary
This dataset combines historical LLM inference prices with benchmark performance scores to construct the largest publicly available benchmark-level LLM price dataset we are aware of. It covers 100+ models across three major benchmarks (GPQA-Diamond, SWE-bench Verified, and AIME) over a two-year window from April 2024 to April 2026, with varying coverage per benchmark.
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-noname/econ_eval.Qwen2.5-Coder-0.5B-Flutter-steps-eval
Qwen2.5-Coder-0.5B Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In steps
mode, the model is given an existing file and an edit instruction and generates a
sequence of localized search/replace edit actions, each mechanically applied to the
current file state before the next action is generated, until the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps-eval.text-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.chess-sft-eval
Chess SFT Eval & Benchmark
Held-out evaluation splits and a frozen benchmark for the
Chess SFT training pipeline.
Every FEN in these files is excluded from training data via a blocklist to guarantee
zero contamination.
Eval examples
13,000
Benchmark examples
13,000
Splits
9 (perception, rules, tactics, evaluation, openings, endgames, planning, chess960, mate)
Format
JSONL
Training companion
Chess-Nut-Engine/chess-sft-data
How eval and benchmark differ… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-eval.swe-race
SWE-Race (public split) v0.1
95 real concurrency and async-lifecycle bugs from 58 open-source Python projects, packaged as
Harbor tasks (the DeepSWE layout) and graded by each project's own hidden
tests in a sealed, offline container. A private split of 93 tasks is held out to detect overfitting.
Leaderboard, per-task results and every agent trajectory: https://labs.evaligo.com/swe-race
Get new results by email: https://labs.evaligo.com/swe-race#signup-form
Full task set… See the full description on the dataset page: https://huggingface.co/datasets/evaligo/swe-race.global-mmlu-rephrased
global_mmlu (rephrased for base-model evaluation)
Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.recon-eval
ReconEval — Financial Reconciliation Benchmark
Reading results from this benchmark. Four properties of ReconEval shape what
a score on it means. Anyone comparing models here should know them.
One class can dominate a margin. PARTIAL_MATCH is the highest-variance class
between models, and its 32 evaluation items are generated from 9 abbreviation
pairs — all of which also appear in the training split, overlap fraction 1.0. On
this set, "learned the concept" and "memorised nine… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/recon-eval.linuxarena-trajectories
LinuxArena Trajectories
Full agent trajectories from LinuxBench/LinuxArena evaluations
across 14 model/policy combinations and 10 environments.
Dataset Description
Each row is one complete evaluation trajectory — every tool call the agent made from start
to finish, with arguments, outputs, errors, and reasoning. Actions are represented as
parallel variable-length lists (one element per action).
Two granularity levels are provided per action:
Normalized… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/linuxarena-trajectories.
