Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hlido-eu /agent-benchmark configs: - config_name: default data_files: - path: run-log.json split: train language: - en license: cc-by-nc-4.0 tags: - ai-agents - benchmark - evaluation - agent-evaluation - llm-agents - trustworthy-ai - product-review - c2pa pretty_name: Hlido AI Agent Benchmark size_categories: - n<1K task_categories: - other Hlido AI Agent Benchmark Independent, cryptographically-attested evaluations of AI agents and products. Hlido reviews AI… See the full description on the dataset page: https://huggingface.co/datasets/hlido-eu/agent-benchmark.1 likes2.4k downloads1d agoHugging Face02vals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K11 likes1.9k downloads1y agoHugging Face03Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads3d agoHugging Face04Intelligent-Internet /ii-agent_gaia-benchmark_validationtextn<1K8 likes830 downloads1y agoHugging Face05obaydata /mcp-agent-trajectory-benchmark MCP Agent Trajectory Benchmark A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces. Designed for training and evaluating tool-use / function-calling capabilities of LLMs. Overview Item Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.texttext-generationn<1K3 likes695 downloads7mo agoHugging Face06RegalFire /Agent-Failure-Recovery-Benchmark Agent Failure Recovery Benchmark 22,573 synthetic, source-verified failure → recovery trajectories across four domains. Evaluate whether a system can reject a failed plan, choose a recovery, or recognize that no recovery exists. Every row links to a hashed record in a pinned public source and is checked by independent computational oracles and trajectory replay. Tasks: failure diagnosis, recovery-action prediction, planning regression, scoped agent evaluation and imitation… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Agent-Failure-Recovery-Benchmark.textreinforcement-learning10K<n<100K1 likes461 downloads4d agoHugging Face07design-agent /in2n-benchmark-bundleimage10K<n<100K0 likes384 downloads5mo agoHugging Face08DataCanvasAILab /Titan-CV-Agent-Benchmark Titan CV Agent Benchmark [中文版] The Titan CV Benchmark is primarily designed to evaluate the performance of AGENTS in the field of computer vision (CV). We have collected more than 200 test examples to comprehensively test the performance of the agent, particularly their capability to solve problems step by step. [!IMPORTANT] Demo Video: https://youtu.be/dcUb4lUnGj4 Github: https://github.com/DataCanvasAILab/Titan-CV-Agent-Benchmark Media files are also available for download via… See the full description on the dataset page: https://huggingface.co/datasets/DataCanvasAILab/Titan-CV-Agent-Benchmark.imagevisual-question-answeringn<1K2 likes361 downloads1y agoHugging Face09Nithish2410 /benchmark-bcplus_agent0 likes329 downloads6mo agoHugging Face10springofwindslabs /mcp-agent-trajectory-benchmark 🛑 Stop LLM Agent Collapse Caused by State Corruption Most agent failures are not tool failures. They are state failures. The model believes the world is still valid — when reality has already changed. This MCP trajectory dataset trains belief revision, state recovery, and autonomous replanning: believed : 45/50 rooms synced ← agent trusts a stale worker log tool : "success", count: 0 ← silent no-op (the worst failure mode) verify : GDS reports actual = 40 ←… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.texttext-generationn<1K0 likes320 downloads2d agoHugging Face11rogue-security /coding-agent-security-benchmark Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.textn<1K2 likes276 downloads2mo agoHugging Face12lyyang2766 /Passing-the-Turing-Test-on-Screen-Agent-Humanization-Benchmark0 likes208 downloads8mo agoHugging Face13obaydata /claude-agent-skills-benchmark Claude Agent Skills Benchmark Claude Agent Skills 评测数据集 Description A benchmark dataset for evaluating whether LLMs can accurately trigger and execute domain-specific Skills on the Claude Code platform. Skills are designed by vertical domain experts with varying complexity levels (based on attachments: scripts, references, assets, and reference markdown files). Evaluation Scenarios Cover: Office automation, coding, investment promotion, financial services, industrial… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/claude-agent-skills-benchmark.documenttext-generationn<1K1 likes198 downloads6mo agoHugging Face14noel7Y /data-agent-benchmarks LongHorizon Full Data-Agent Benchmarks Companion data artifacts for five complete evaluation tracks: DataSciBench full55 / 167 metric entries DABStep full450 DABStep-Research full100 DSBench Modeling full74 LongDS full68 / 2,225 turns The companion GitHub repository contains processed manifests, evaluation code, historical API ReAct baseline code, download/preparation tools, and the frozen source lock. artifact_manifest.json records every uploaded object's size, SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.imagequestion-answering1 likes178 downloads1mo agoHugging Face15willykeenan /agent-handoff-benchmark Agent Handoff Benchmark A text-free evaluation dataset for studying whether partial classifier output preserves a downstream decision. It packages the public Waggle/Kea BANKING77 reference artifacts as three directly loadable tables: intent predictions, 77-class probability vectors, and decision/certificate outcomes under 12 fixed synthetic cost policies. William Keenan created the classifier experiment, derived predictions, decision-sufficiency experiment and this packaging.… See the full description on the dataset page: https://huggingface.co/datasets/willykeenan/agent-handoff-benchmark.tabular10K<n<100K2 likes174 downloads13d agoHugging Face16sumitaidev /agent-memory-resilience-benchmark Agent Memory Resilience & Poisoning Benchmark Dataset Summary This benchmark dataset evaluates resilience, negative transfer, and memory poisoning mitigation in autonomous LLM agent architectures (such as LangGraph, AutoGen, and CrewAI). When autonomous agents record distilled self-reflections after attempting tasks, external stochastic failures or subtle API deprecations often cause agents to commit defective strategies into episodic memory. Under standard… See the full description on the dataset page: https://huggingface.co/datasets/sumitaidev/agent-memory-resilience-benchmark.tabularreinforcement-learning1K<n<10K1 likes137 downloads17d agoHugging Face17Omcrec /ecommerce-ai-data-analyst-agent-benchmark E-commerce AI Data Analyst Agent Benchmark A synthetic e-commerce dataset for evaluating AI data analyst agents on realistic, multi-step business analysis, data-quality investigation, and analytical reasoning. This dataset is part of the E-commerce AI Data Analyst Agent Benchmark. Dataset summary This dataset supports evaluation of AI data analyst agents on realistic, multi-step e-commerce analysis. It contains: customers.csv products.csv orders.csv returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.tabulartable-question-answering1 likes132 downloads25d agoHugging Face18mcphunt-benchmark /mcphunt-agent-traces MCPHunt Agent Traces Agent execution traces from the MCPHunt evaluation framework, measuring cross-boundary data propagation in multi-server MCP agents. Contents main/ — 3,615 traces from 5 models across 147 tasks and 7 environment variants (risky_v1/v2/v3, benign, hard_neg_v1/v2/v3). One JSON file per model. mitigation/ — 2,706 traces from the prompt-mitigation study (M0--M3 levels) across 3 models. live_guard_defense/ — 387 DeepSeek-V4-Flash traces from the… See the full description on the dataset page: https://huggingface.co/datasets/mcphunt-benchmark/mcphunt-agent-traces.tabularothern<1K3 likes129 downloads2mo agoHugging Face19Koplos /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/finance_agent_benchmark.textn<1K1 likes126 downloads2mo agoHugging Face20anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes117 downloads3mo agoHugging Face21chatjesus /voice-agent-benchmark-landscape Voice Agent Benchmark Landscape A structured, source-linked map of public benchmarks for voice agents, spoken assistants, speech-enabled tool use, computer action, ASR, and meeting understanding. This is a landscape dataset, not a leaderboard. Each row records whether a benchmark covers spoken input/output, multi-turn interaction, tools, goal completion, computer or browser action, meeting or long-form content, real-time operation, and public data. Values are deliberately… See the full description on the dataset page: https://huggingface.co/datasets/chatjesus/voice-agent-benchmark-landscape.textn<1K3 likes114 downloads1mo agoHugging Face22Agent-Threat-Rule /atr-skill-benchmark ATR Skill-Security Benchmark A labeled corpus of SKILL.md files for evaluating detection of malicious agent skills — prompt injection, tool poisoning, credential theft, malware droppers and supply-chain attacks hidden inside natural-language agent instructions. Published as part of Agent Threat Rules (ATR), an open, vendor-neutral detection standard for AI agents (like Sigma, but for agent attacks). Why this exists SKILL.md files are natural-language instructions… See the full description on the dataset page: https://huggingface.co/datasets/Agent-Threat-Rule/atr-skill-benchmark.texttext-classificationn<1K2 likes94 downloads3mo agoHugging Face23KikoCis /real-world-agent-benchmark Real-World Agent Benchmark (RAB) Paper: Orchestrator and Task-Framing Effects Dominate Fine-Tuning in Real-World Agent Evaluation of a Quantized 31B ModelAuthors: Kiko Cisneros, Claude Sonnet 4.6 · Utopia IA, May 2026Code: github.com/KikoCisBot/gemma4-31b-study 📄 See paper4_orchestrator_dominance.pdf in the Files tab. TL;DR Standard benchmarks (BFCL, HumanEval) do not predict real-world agent capability. A model scoring 95% BFCL scores 0/10 on a real autonomous task… See the full description on the dataset page: https://huggingface.co/datasets/KikoCis/real-world-agent-benchmark.documentn<1K1 likes83 downloads5mo agoHugging Face24kushalicious /agent-memory-benchmark Agent Memory Compression & Evaluation Benchmark This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics. Dataset Structure 1. conversation.json A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under: Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.textquestion-answeringn<1K0 likes81 downloads4mo agoHugging Face25clicksprotocol /agent-treasury-benchmark Agent Treasury Benchmark Simulated treasury performance data for AI agents using Clicks Protocol on Base. What is this? A benchmark dataset for evaluating agent treasury strategies: how AI agents manage idle USDC through automated yield protocols. Each row represents one agent's monthly treasury snapshot. Use this to: Benchmark different yield allocation strategies (5-50% yield split) Model agent economics under varying DeFi conditions Train models to predict optimal… See the full description on the dataset page: https://huggingface.co/datasets/clicksprotocol/agent-treasury-benchmark.tabular-regression1K<n<10K0 likes74 downloads6mo agoHugging Face26ruchit11111 /coding-agent-security-benchmark Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.textn<1K1 likes74 downloads1mo agoHugging Face27razi06608333 /ai-agent-failure-safety-benchmark-free-sample1 likes73 downloads18d agoHugging Face28akseljoonas /hf-agent-benchmarktextn<1K0 likes69 downloads1y agoHugging Face29amitmaity0 /local-coding-agent-benchmark 🤖 Local Coding Agent Benchmark (LCAB) Real-world benchmarking of local AI coding agents on software-repair workloads. This Hugging Face Dataset contains the reproducibility artifacts, raw agent-session evidence, benchmark results, task source, hardware profiles, and analysis for the Local Coding Agent Benchmark (LCAB). LCAB is designed to evaluate local coding agents as complete systems—not only by tokens/second, but by how efficiently they transform a real software-repair… See the full description on the dataset page: https://huggingface.co/datasets/amitmaity0/local-coding-agent-benchmark.2 likes66 downloads2mo agoHugging Face30NotoriousH2 /calendar-agent-benchmark Calendar Agent Training and Evaluation Tool Calling SFT와 환경 기반 RLVR을 함께 실습하는 Calendar Agent 데이터입니다. 1. 데이터 구성 파일 개수 구성 SHA-256 sft_train.jsonl 3,000 기본 시나리오, 다중 참석자 생성, 안전 대조군 f882fba3a6c32982d2f8cca06465fa1ec3c68562c655d2ee1bcd92decd29e637 sft_validation.jsonl 300 SFT 학습 중 검증 b657e4162ef26bd580ad975b4f2af575a616dda17c90c7b4edab38790ad66ac5 rlvr_train.jsonl 256 충돌 복구와 다중 참석자 생성 128개, 안전 경계 128개… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/calendar-agent-benchmark.text-generation1 likes58 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.