Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agents-course /course-imagesimagen<1K21 likes330k downloads1y agoHugging Face02mercor /apex-agentsgated APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents.documentn<1K203 likes99k downloads4mo agoHugging Face03agents-last-exam /agents-last-exam-data Agents Last Exam — Task Input Data Input files (the materials each task hands to the agent at run start) for the Agents Last Exam (ALE) benchmark. Browsable per-task directory layout. The Agents Last Exam dataset family ALE is published as three companion HuggingFace datasets: Dataset Contents Access Task Card Metadata One row per task: titles, prompts, taxonomy, input-file descriptors Open Task Input Data The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data.7 likes40k downloads3d agoHugging Face04agentica-org /DeepScaleR-Preview-Dataset Data Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from: AIME (American Invitational Mathematics Examination) problems (1984-2023) AMC (American Mathematics Competition) problems (prior to 2023) Omni-MATH dataset Still dataset Format Each row in the JSON dataset contains: problem: The mathematical question text, formatted with LaTeX notation. solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.text10K<n<100K208 likes39k downloads2y agoHugging Face05meta-agents-research-environments /gaia2 Gaia2 Paper | Code | Project Page Dataset Summary Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically. The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.textreinforcement-learningn<1K47 likes38k downloads1y agoHugging Face06Autonomous-Scientific-Agents /results8 likes38k downloads2mo agoHugging Face07openbmb /UltraData-SFT-Agent-2609 UltraData-SFT-Agent-2609 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series English | 中文 📚 Introduction UltraData-SFT-Agent-2609 is the L3 refined data for Agent instruction-tuning within UltraData's L0-L4 tiered data management framework. Built for the post-training of MiniCPM5-2B, it complements UltraData-SFT-2605 (core-domain SFT) with executable Agent trajectories. The release contains approximately 500,000 samples spanning tool use… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609.texttext-generation100K<n<1M292 likes27k downloads1mo agoHugging Face08McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes23k downloads1y agoHugging Face09agents-last-exam /ale-images-qcow2 ALE QEMU runner image agentslastexam/ale-qemu is the container-side runtime used by the ALE qemu provider. It packages QEMU, KVM integration, NAT networking, noVNC, and process supervision. The Ubuntu or Windows guest is supplied separately as /storage/data.qcow2. Docker is the container runtime. Dockur is the upstream QEMU-in-Docker project whose startup and networking stack this image inherits. ALE adds a stable runner contract around that upstream image. The image is based on… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/ale-images-qcow2.0 likes22k downloads2d agoHugging Face10agents-course /unit4-students-scores20 likes19k downloads19m agoHugging Face11TMaxxx /agent-task-recursive-task-synthesis Apptainer pool for hamishivi/agent-task-recursive-task-synthesis This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-recursive-task-synthesis. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image. Apptainer images The pool currently contains 29,501 / 29… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-recursive-task-synthesis.5 likes16k downloads22d agoHugging Face12agents-course /certificates97 likes15k downloads3m agoHugging Face13neulab /agent-data-collection Agent Data Collection A comprehensive collection of agent interaction datasets for training and evaluating AI agents across diverse domains and tasks. This dataset aggregates high-quality agent trajectories from various environments including web browsing, code generation, household tasks, knowledge base querying, and software engineering. The dataset is collected through methods described in Agent Data Protocol. Dataset Splits Each dataset configuration provides up… See the full description on the dataset page: https://huggingface.co/datasets/neulab/agent-data-collection.text-generation1M<n<10M116 likes11k downloads7mo agoHugging Face14meta-agents-research-environments /gaia2_filesystem GAIA2 Filesystem This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset. Dataset Link https://huggingface.co/datasets/meta-agents-research-environments/gaia2 Contact Details Publishing POC: Meta AI Research Team Affiliation: Meta Platforms, Inc. Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.imagen<1K1 likes11k downloads1y agoHugging Face15Autonomous-Scientific-Agents /requests7 likes10k downloads2mo agoHugging Face16agents-last-exam /agents-last-exam-referencegated Agents Last Exam — Reference (Ground-Truth) Data ⚠️ Gated dataset. This repo contains the ground-truth / reference outputs used to score the Agents Last Exam (ALE) benchmark. Access requires login, agreement to the terms on the access-request form, and manual approval. Note (06/16/26): This repository was accidentally deleted and has been recreated. The previous list of approved requesters could not be restored, so even if you were granted access before, you will need to… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-reference.9 likes8.8k downloads3d agoHugging Face17skeole /qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols. ~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks. The only human artifacts are: agents/* human/* AGENTS.md texttext-generation1K<n<10K3 likes8.7k downloads20d agoHugging Face18HanXiao1999 /UI-Genie-Agent-16kThis repository contains the Trajectory dataset from the paper UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents. Github: https://github.com/Euphoria16/UI-Genie imageimage-text-to-text10K<n<100K0 likes8.7k downloads11mo agoHugging Face19nebius /SWE-agent-trajectories Dataset Summary This dataset contains 80,036 trajectories generated by a software engineering agent based on the SWE-agent framework, using various models as action generators. In these trajectories, the agent attempts to solve GitHub issues from the nebius/SWE-bench-extra and the dev split of princeton-nlp/SWE-bench. Dataset Description This dataset was created as part of a research project focused on developing a software engineering agent using open-weight models… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-agent-trajectories.text10K<n<100K100 likes8.5k downloads2y agoHugging Face20nvidia /Nemotron-SFT-Agentic-v2 Dataset Description The Nemotron-SFT-Agentic-v2 dataset is a collection of synthetic single-turn and multi-turn tool-use trajectories designed to strengthen models’ capabilities as interactive, tool-using agents. It targets tasks where the model must decompose user goals, decide when to call tools, and reason over tool outputs to complete tasks reliably and safely. This dataset is ready for commercial use. The dataset consolidates three internally curated components (described… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Agentic-v2.text-generation86 likes8.4k downloads2mo agoHugging Face21Stage-jh-monitor /appworld-qwen35-4b-agent-rl-epoch3 appworld-qwen35-4b-agent-rl-epoch3 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.45859375 Action score: 0.475 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face22Stage-jh-monitor /appworld-qwen35-4b-agent-rl-epoch3-reeval1 appworld-qwen35-4b-agent-rl-epoch3-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4578125 Action score: 0.4921875 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face23ai-safety-institute /AgentHarm AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Maksym Andriushchenko1,†,*, Alexandra Souly2,* Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,* Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,* 1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor †EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.textn<1K65 likes8.3k downloads2y agoHugging Face24Agent-Ark /Toucan-1.5M 🦤 Toucan-1.5M: Toucan-1.5M is the largest fully synthetic tool-agent dataset to date, designed to advance tool use in agentic LLMs. It comprises over 1.5 million trajectories synthesized from 495 real-world Model Context Protocols (MCPs) spanning 2,000+ tools. By leveraging authentic MCP environments, Toucan-1.5M generates diverse, realistic, and challenging tasks requires using multiple tools, with trajectories involving real tool executions across multi-round, multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Agent-Ark/Toucan-1.5M.text1M<n<10M238 likes8.1k downloads1y agoHugging Face25xlangai /AgentNet OpenCUA: Open Foundations for Computer-Use Agents 🌐 Website 🔎 Data Viewer 📝 Paper 💻 Code AgentNet Dataset AgentNet is the first large-scale desktop computer-use agent trajectory dataset, containing 22.6K human-annotated computer-use tasks across Windows, macOS, and Ubuntu systems. Applications This dataset enables training and evaluation of: Vision-language-action (VLA) models for computer use Multi-modal agents… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/AgentNet.image-text-to-text95 likes7.4k downloads9mo agoHugging Face26ICML-2026-agent-repro /verdicts Logbook verdicts Per-logbook claim verdicts produced by the logbook-judge Space. See verdicts.json. 3 likes7.1k downloads2mo agoHugging Face27TMaxxx /agent-task-facet-terminal-6k Apptainer pool for hamishivi/agent-task-facet-terminal-6k This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-facet-terminal-6k. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image. Apptainer images The pool currently contains 6,020 / 6,020 verified… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-facet-terminal-6k.4 likes6.8k downloads23d agoHugging Face28TMaxxx /agent-task-terminal-lego-15k Apptainer pool for hamishivi/agent-task-terminal-lego-15k This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-terminal-lego-15k. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image. Apptainer images The pool currently contains 15,048 / 15,048 verified… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-terminal-lego-15k.2 likes6.6k downloads23d agoHugging Face29AgentNativeResearchLab /oeb-scored-runs Open-Endedness Bench: scored runs Every agent run scored in the paper Open-Endedness Bench: Measuring Epistemic Process from Agent Records, with the output of each scoring stage and the full log of judge requests and answers. The code is at github.com/ARA-Labs/oeb. Layout out/posttrainbench/<task panel>/<unit>/ PostTrainBench runs (post-training gemma-3-4b on six held-out tasks); a model's second run is <model>-r2… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/oeb-scored-runs.0 likes6.4k downloads11d agoHugging Face30AgentNativeResearchLab /arc-agi3-codex-gpt5.5-su15 ARC-AGI-3 su15 — Agent Trajectories (codex-gpt5.5) Gameplay trajectories from the harness×model pair codex-gpt5.5 playing the ARC-AGI-3 game su15, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.5-su15.reinforcement-learning0 likes6.3k downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.