datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-benchmark
configs:
- config_name: default
data_files:
- path: run-log.json
split: train
language:
- en
license: cc-by-nc-4.0
tags:
- ai-agents
- benchmark
- evaluation
- agent-evaluation
- llm-agents
- trustworthy-ai
- product-review
- c2pa
pretty_name: Hlido AI Agent Benchmark
size_categories:
- n<1K
task_categories:
- other
Hlido AI Agent Benchmark
Independent, cryptographically-attested evaluations of AI agents and products.
Hlido reviews AI… See the full description on the dataset page: https://huggingface.co/datasets/hlido-eu/agent-benchmark.finance_agent_benchmark
Finance Agent Benchmark Dataset
We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings.
We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.ii-agent_gaia-benchmark_validationmcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.Agent-Failure-Recovery-Benchmark
Agent Failure Recovery Benchmark
22,573 synthetic, source-verified failure → recovery trajectories across four domains.
Evaluate whether a system can reject a failed plan, choose a recovery, or recognize that no recovery exists. Every row links to a hashed record in a pinned public source and is checked by independent computational oracles and trajectory replay.
Tasks: failure diagnosis, recovery-action prediction, planning regression, scoped agent evaluation and imitation… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Agent-Failure-Recovery-Benchmark.in2n-benchmark-bundleTitan-CV-Agent-Benchmark
Titan CV Agent Benchmark
[中文版]
The Titan CV Benchmark is primarily designed to evaluate the performance of AGENTS in the field of computer vision (CV). We have collected more than 200 test examples to comprehensively test the performance of the agent, particularly their capability to solve problems step by step.
[!IMPORTANT]
Demo Video: https://youtu.be/dcUb4lUnGj4
Github: https://github.com/DataCanvasAILab/Titan-CV-Agent-Benchmark
Media files are also available for download via… See the full description on the dataset page: https://huggingface.co/datasets/DataCanvasAILab/Titan-CV-Agent-Benchmark.benchmark-bcplus_agentmcp-agent-trajectory-benchmark
🛑 Stop LLM Agent Collapse Caused by State Corruption
Most agent failures are not tool failures. They are state failures.
The model believes the world is still valid — when reality has already changed.
This MCP trajectory dataset trains belief revision, state recovery, and autonomous replanning:
believed : 45/50 rooms synced ← agent trusts a stale worker log
tool : "success", count: 0 ← silent no-op (the worst failure mode)
verify : GDS reports actual = 40 ←… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.Passing-the-Turing-Test-on-Screen-Agent-Humanization-Benchmarkclaude-agent-skills-benchmark
Claude Agent Skills Benchmark
Claude Agent Skills 评测数据集
Description
A benchmark dataset for evaluating whether LLMs can accurately trigger and execute domain-specific Skills on the Claude Code platform. Skills are designed by vertical domain experts with varying complexity levels (based on attachments: scripts, references, assets, and reference markdown files).
Evaluation Scenarios Cover:
Office automation, coding, investment promotion, financial services, industrial… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/claude-agent-skills-benchmark.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.agent-handoff-benchmark
Agent Handoff Benchmark
A text-free evaluation dataset for studying whether partial classifier output
preserves a downstream decision. It packages the public Waggle/Kea BANKING77
reference artifacts as three directly loadable tables: intent predictions,
77-class probability vectors, and decision/certificate outcomes under 12 fixed
synthetic cost policies.
William Keenan created the classifier experiment, derived predictions,
decision-sufficiency experiment and this packaging.… See the full description on the dataset page: https://huggingface.co/datasets/willykeenan/agent-handoff-benchmark.agent-memory-resilience-benchmark
Agent Memory Resilience & Poisoning Benchmark
Dataset Summary
This benchmark dataset evaluates resilience, negative transfer, and memory poisoning mitigation in autonomous LLM agent architectures (such as LangGraph, AutoGen, and CrewAI).
When autonomous agents record distilled self-reflections after attempting tasks, external stochastic failures or subtle API deprecations often cause agents to commit defective strategies into episodic memory. Under standard… See the full description on the dataset page: https://huggingface.co/datasets/sumitaidev/agent-memory-resilience-benchmark.ecommerce-ai-data-analyst-agent-benchmark
E-commerce AI Data Analyst Agent Benchmark
A synthetic e-commerce dataset for evaluating AI data analyst agents on
realistic, multi-step business analysis, data-quality investigation, and
analytical reasoning.
This dataset is part of the
E-commerce AI Data Analyst Agent Benchmark.
Dataset summary
This dataset supports evaluation of AI data analyst agents on realistic,
multi-step e-commerce analysis.
It contains:
customers.csv
products.csv
orders.csv
returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.mcphunt-agent-traces
MCPHunt Agent Traces
Agent execution traces from the MCPHunt evaluation framework, measuring
cross-boundary data propagation in multi-server MCP agents.
Contents
main/ — 3,615 traces from 5 models across 147 tasks and 7 environment
variants (risky_v1/v2/v3, benign, hard_neg_v1/v2/v3). One JSON file per model.
mitigation/ — 2,706 traces from the prompt-mitigation study (M0--M3
levels) across 3 models.
live_guard_defense/ — 387 DeepSeek-V4-Flash traces from the… See the full description on the dataset page: https://huggingface.co/datasets/mcphunt-benchmark/mcphunt-agent-traces.finance_agent_benchmark
Finance Agent Benchmark Dataset
We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings.
We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/finance_agent_benchmark.structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.voice-agent-benchmark-landscape
Voice Agent Benchmark Landscape
A structured, source-linked map of public benchmarks for voice agents, spoken assistants, speech-enabled tool use, computer action, ASR, and meeting understanding.
This is a landscape dataset, not a leaderboard. Each row records whether a benchmark covers spoken input/output, multi-turn interaction, tools, goal completion, computer or browser action, meeting or long-form content, real-time operation, and public data.
Values are deliberately… See the full description on the dataset page: https://huggingface.co/datasets/chatjesus/voice-agent-benchmark-landscape.atr-skill-benchmark
ATR Skill-Security Benchmark
A labeled corpus of SKILL.md files for evaluating detection of malicious agent
skills — prompt injection, tool poisoning, credential theft, malware droppers
and supply-chain attacks hidden inside natural-language agent instructions.
Published as part of Agent Threat Rules (ATR),
an open, vendor-neutral detection standard for AI agents (like Sigma, but for
agent attacks).
Why this exists
SKILL.md files are natural-language instructions… See the full description on the dataset page: https://huggingface.co/datasets/Agent-Threat-Rule/atr-skill-benchmark.real-world-agent-benchmark
Real-World Agent Benchmark (RAB)
Paper: Orchestrator and Task-Framing Effects Dominate Fine-Tuning in Real-World Agent Evaluation of a Quantized 31B ModelAuthors: Kiko Cisneros, Claude Sonnet 4.6 · Utopia IA, May 2026Code: github.com/KikoCisBot/gemma4-31b-study
📄 See paper4_orchestrator_dominance.pdf in the Files tab.
TL;DR
Standard benchmarks (BFCL, HumanEval) do not predict real-world agent capability. A model scoring 95% BFCL scores 0/10 on a real autonomous task… See the full description on the dataset page: https://huggingface.co/datasets/KikoCis/real-world-agent-benchmark.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.agent-treasury-benchmark
Agent Treasury Benchmark
Simulated treasury performance data for AI agents using Clicks Protocol on Base.
What is this?
A benchmark dataset for evaluating agent treasury strategies: how AI agents manage idle USDC through automated yield protocols. Each row represents one agent's monthly treasury snapshot.
Use this to:
Benchmark different yield allocation strategies (5-50% yield split)
Model agent economics under varying DeFi conditions
Train models to predict optimal… See the full description on the dataset page: https://huggingface.co/datasets/clicksprotocol/agent-treasury-benchmark.coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.ai-agent-failure-safety-benchmark-free-samplehf-agent-benchmarklocal-coding-agent-benchmark
🤖 Local Coding Agent Benchmark (LCAB)
Real-world benchmarking of local AI coding agents on software-repair workloads.
This Hugging Face Dataset contains the reproducibility artifacts, raw agent-session evidence, benchmark results, task source, hardware profiles, and analysis for the Local Coding Agent Benchmark (LCAB).
LCAB is designed to evaluate local coding agents as complete systems—not only by tokens/second, but by how efficiently they transform a real software-repair… See the full description on the dataset page: https://huggingface.co/datasets/amitmaity0/local-coding-agent-benchmark.calendar-agent-benchmark
Calendar Agent Training and Evaluation
Tool Calling SFT와 환경 기반 RLVR을 함께 실습하는 Calendar Agent 데이터입니다.
1. 데이터 구성
파일
개수
구성
SHA-256
sft_train.jsonl
3,000
기본 시나리오, 다중 참석자 생성, 안전 대조군
f882fba3a6c32982d2f8cca06465fa1ec3c68562c655d2ee1bcd92decd29e637
sft_validation.jsonl
300
SFT 학습 중 검증
b657e4162ef26bd580ad975b4f2af575a616dda17c90c7b4edab38790ad66ac5
rlvr_train.jsonl
256
충돌 복구와 다중 참석자 생성 128개, 안전 경계 128개… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/calendar-agent-benchmark.
