Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mercor /apex-agentsgated APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents.documentn<1K203 likes99k downloads4mo agoHugging Face02meta-agents-research-environments /gaia2 Gaia2 Paper | Code | Project Page Dataset Summary Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically. The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.textreinforcement-learningn<1K47 likes38k downloads1y agoHugging Face03agents-course /unit4-students-scorestext10K<n<100K20 likes19k downloads2h agoHugging Face04meta-agents-research-environments /gaia2_filesystem GAIA2 Filesystem This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset. Dataset Link https://huggingface.co/datasets/meta-agents-research-environments/gaia2 Contact Details Publishing POC: Meta AI Research Team Affiliation: Meta Platforms, Inc. Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.imagen<1K1 likes11k downloads1y agoHugging Face05agents-last-exam /agents-last-exam Agents Last Exam — Task Card Metadata (v1.1) A metadata-only release (v1.1) of 151 tasks from the Agents Last Exam (ALE) benchmark for evaluating computer-use agents on long-horizon professional work. The Agents Last Exam dataset family ALE is published as three companion HuggingFace datasets: Dataset Contents Access Task Card Metadata One row per task: titles, prompts, taxonomy, input-file descriptors Open Task Input Data The input/ files each task… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam.textn<1K214 likes3.8k downloads2d agoHugging Face06SciPhi /AgentSearch-V1 Getting Started The AgentSearch-V1 dataset boasts a comprehensive collection of over one billion embeddings, produced using jina-v2-base. The dataset encompasses more than 50 million high-quality documents and over 1 billion passages, covering a vast range of content from sources such as Arxiv, Wikipedia, Project Gutenberg, and includes carefully filtered Creative Commons (CC) data. Our team is dedicated to continuously expanding and enhancing this corpus to improve the search… See the full description on the dataset page: https://huggingface.co/datasets/SciPhi/AgentSearch-V1.texttext-generation10K<n<100K92 likes3.7k downloads3y agoHugging Face07mercor /apex-agents-v1.1gated APEX-Agents 1.1 APEX-Agents 1.1 is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional-services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers. They require agents to work across realistic project files and applications such as documents, spreadsheets, PDFs, email, chat, and calendar. Tasks: 240 total (80 per job category) Worlds: 31 total (8 investment… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents-v1.1.textn<1K8 likes3k downloads4d agoHugging Face08NexusProjectsAI /Nexus-Agents-ToolCalling Nexus Agents — Tool-Calling Conversations Synthetic, schema-verified tool-calling conversations for training the Nexus Projects agents. This is the exact data behind Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF), including the verification transcripts that scored it (27/27 on the behavioral interview eval, vs 13/27 for the base model). Links: the fine-tuned model → Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) · the generator + seed data + eval harness → Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.texttext-generation100K<n<1M2 likes2.4k downloads4mo agoHugging Face09amanutej /trustworthy-biology-agents-traces Trustworthy Biology Agents — Run Traces Raw execution traces from 1,329 agent runs across three coding agents on three biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed trace bundle for the study in manu-tej/ai-scientists; the write-up lives in that repo's RESULTS.md. The motivating question is not only whether an agent reaches the right answer, but whether it behaves like a trustworthy analyst when the task is ambiguous, under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.tabular1K<n<10K0 likes2.2k downloads3mo agoHugging Face10agents-course /course-certificates-of-excellencetext1K<n<10K14 likes1.6k downloads2h agoHugging Face11agentsea /wave-uiLICENSE image10K<n<100K27 likes1.3k downloads2y agoHugging Face12Lyric1010 /agent-sft-10B Dataset: agent-sft-10B This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/agent-sft-10B/no-curriculum/tmp. text0 likes1.1k downloads11mo agoHugging Face13agents-course /unit3-inviteestextn<1K21 likes992 downloads2y agoHugging Face14lobinni /apex-agents APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12 law)… See the full description on the dataset page: https://huggingface.co/datasets/lobinni/apex-agents.documentn<1K0 likes783 downloads8mo agoHugging Face15idleengine /apex-agents APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/apex-agents.documentn<1K0 likes763 downloads2mo agoHugging Face16agentsea /wave-ui-25k WaveUI-25k This dataset contains 25k examples of labeled UI elements. It is a subset of a collection of ~80k preprocessed examples assembled from the following sources: WebUI RoboFlow GroundUI-18K These datasets were preprocessed to have matching schemas and to filter out unwanted examples, such as duplicated, overlapping and low-quality datapoints. We also filtered out many text elements which were not in the main scope of this work. The WaveUI-25k dataset includes the original… See the full description on the dataset page: https://huggingface.co/datasets/agentsea/wave-ui-25k.image10K<n<100K40 likes741 downloads2y agoHugging Face17ZixuanKe /evovling_agents Evolving Agents Benchmark This repository contains the data presented in EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?. Project page: https://mas-orchestra.salesforceresearch.ai/evoharness/ A versioned, per-split, multi-domain library of given Codex subagents, produced by evovle_agents. It is the agent-track analogue of evovling_tools: where evovling_skills evaluates a model that generates skills, evolving-agents evaluates a model that orchestrates given… See the full description on the dataset page: https://huggingface.co/datasets/ZixuanKe/evovling_agents.textother1K<n<10K0 likes707 downloads1mo agoHugging Face18secondstate /finance-agents-benchmark-traces FAB — Agent Traces and Grading 600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included. The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.… See the full description on the dataset page: https://huggingface.co/datasets/secondstate/finance-agents-benchmark-traces.tabularquestion-answeringn<1K1 likes678 downloads13d agoHugging Face19Jaward /lectura-agents-data LectūraAgents Dataset Overview This dataset is in support of findings in our paper LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching. LectūaAgents is a hierarchical multi-agent framework that enables end-to-end personalized learning experiences through adaptive embodied teaching. It mirrors a professor–students’ relationship, wherein a ProfessorAgent guides a collaborative team of specialized subordinate… See the full description on the dataset page: https://huggingface.co/datasets/Jaward/lectura-agents-data.audion<1K24 likes610 downloads1mo agoHugging Face20voidful /agent-sft-stitch-zh-tts agent-sft-stitch-zh-tts Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted. Configs records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.audiotext-to-speech100K<n<1M0 likes592 downloads3mo agoHugging Face21BAAI-Agents /SWITCH SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios [arXiv] [leaderboard] [dataset] [PDF] Dataset Summary SWITCH (Semantic World Interface Tasks for Control & Handling) is a multimodal embodied-interaction benchmark for understanding, modeling, and evaluating actions over Tangible Control Interfaces (TCIs) in egocentric real-world scenarios. TCIs include everyday interfaces such as appliance panels, lighting… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Agents/SWITCH.imagevisual-question-answering1K<n<10K6 likes591 downloads1mo agoHugging Face22while-ai /agent-simulations Agent Simulations Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets 53,971 synthetic agent trajectories generated by simulations across 34 agent types. The rows include successful and failed trajectories for supervised fine-tuning, preference work, reinforcement learning, and evaluation. NOTE: This is generated test and training data, not curated ground truth. Review and filter it for your application before training or… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.texttext-generation10K<n<100K0 likes580 downloads18d agoHugging Face23secondstate /finance-agents-benchmark FAB — Finance Agents Benchmark FAB is an open-source project for benchmarking LLM agents' ability to perform financial due diligence in a synthetic company data room. FAB consists of a dataset of tasks containing agent instructions, documents and rubrics, and an execution harness for running and evaluating agents. This repository contains the dataset; the harness is available on GitHub. Dataset 50 tasks · 160 documents · 231 grading criteria · One shared data room… See the full description on the dataset page: https://huggingface.co/datasets/secondstate/finance-agents-benchmark.documentquestion-answeringn<1K1 likes567 downloads13d agoHugging Face24AgentCrush /agents-index AgentCrush Agent Index Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz. Overview 1,450 agents indexed across categories: developer tools, tokenized agents, service agents, model families 221 evidence-ranked with verified multi-signal scores Updated: 2026-10-10 Configs Config Description Rows agents All indexed agents with metadata ~1,450 evidence_ranked Evidence-ranked tier only ~221 snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.tabulartext-classification1K<n<10K0 likes503 downloads12h agoHugging Face25neyralabs /apex-agents APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12 law)… See the full description on the dataset page: https://huggingface.co/datasets/neyralabs/apex-agents.documentn<1K0 likes450 downloads7mo agoHugging Face26meta-agents-research-environments /gaia2-cli GAIA2 CLI Benchmark dataset for gaia2-cli, the CLI-based agent evaluation harness. Schema Each row has two columns: Column Type Description scenario_id string Unique scenario identifier (e.g. scenario_universe_21_1qgjj6) scenario string Complete scenario as a JSON string Usage from datasets import load_dataset import json # Load a specific config (160 scenarios) ds = load_dataset("meta-agents-research-environments/gaia2-cli", "adaptability"… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2-cli.text1K<n<10K0 likes444 downloads6mo agoHugging Face27Self-Improving-Coding-Agents /SI2CA-Training-TrajectoriesDataset Card for SI2CA-Training-Trajectories [🌐 Website] • [🤗 Dataset] • [📜 Paper] • [🐱 GitHub] 💡 Introduction This dataset consists of 32,340 coding-agent trajectories generated by Qwen3.5-122B-A10B on the same 10,780 executable Python SWE tasks under the three trajectory-curation settings of Section 4.4 of the paper: standard sampling, full self-judgement, and an efficient discovered strategy found by the recursive self-improvement framework. Each task is… See the full description on the dataset page: https://huggingface.co/datasets/Self-Improving-Coding-Agents/SI2CA-Training-Trajectories.tabulartext-generation10K<n<100K0 likes405 downloads19d agoHugging Face28voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes380 downloads3mo agoHugging Face29abridges /apex-agents APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12 law)… See the full description on the dataset page: https://huggingface.co/datasets/abridges/apex-agents.documentn<1K0 likes378 downloads7mo agoHugging Face30Agents-X /TIR-Bench TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning Introduction: TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.imagequestion-answering1K<n<10K3 likes353 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.