Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agents-course /course-imagesimagen<1K21 likes330k downloads1y agoHugging Face02mercor /apex-agentsgated APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents.documentn<1K203 likes99k downloads4mo agoHugging Face03McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes23k downloads1y agoHugging Face04meta-agents-research-environments /gaia2_filesystem GAIA2 Filesystem This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset. Dataset Link https://huggingface.co/datasets/meta-agents-research-environments/gaia2 Contact Details Publishing POC: Meta AI Research Team Affiliation: Meta Platforms, Inc. Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.imagen<1K1 likes11k downloads1y agoHugging Face05HanXiao1999 /UI-Genie-Agent-16kThis repository contains the Trajectory dataset from the paper UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents. Github: https://github.com/Euphoria16/UI-Genie imageimage-text-to-text10K<n<100K0 likes8.7k downloads11mo agoHugging Face06yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M53 likes4.2k downloads7mo agoHugging Face07zr-wang /AgenticOCR-SFT AgenticOCR SFT Training Data Supervised fine-tuning data for the AgenticOCR project. The dataset contains 7,631 training records in sft_combined_0422.json. Image paths in each record are relative to the repository root and point into sft_images/. imagevisual-question-answering10K<n<100K2 likes3.5k downloads2mo agoHugging Face08mlfoundations /gelato-osworld-agent-trajectoriesimage10K<n<100K2 likes3.3k downloads11mo agoHugging Face09Yushi123 /Gui-agent Gui-Agent — GUI trajectories in LIBERO/VLA format Human GUI demonstrations from four sources, unified into a single VLA-style intermediate representation and written as LIBERO-layout HDF5, so LIBERO/VLA dataloaders run against GUI data unchanged. raw source ──[adapter]──> GuiEpisode ──[writer]──> LIBERO-style HDF5 per-source the IR format- what you train on only specific 25,872 episodes / 453,264 steps / 235 GB… See the full description on the dataset page: https://huggingface.co/datasets/Yushi123/Gui-agent.imageroboticsn<1K1 likes2.3k downloads2mo agoHugging Face10RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes2.2k downloads7mo agoHugging Face11xbench /AgentIF-OneDay Dataset Card for AgentIF-OneDay Dataset Details AgentIF-OneDay is a comprehensive benchmark designed to evaluate AI agents on diverse, daily tasks across work, life, and learning scenarios. Unlike evaluations focused solely on task difficulty, this dataset emphasizes the breadth of general user needs, requiring agents to handle complex attachments, infer implicit instructions, and deliver tangible file-based outputs. It comprises 104 tasks structured around Open… See the full description on the dataset page: https://huggingface.co/datasets/xbench/AgentIF-OneDay.documentn<1K4 likes2.1k downloads8mo agoHugging Face12Agent-as-Policy /agent-as-policy Agent as Policy — Real Dual-Arm LLM-Agent Manipulation Trials Paper · Project page · Code Summary 162 real-robot trials in which an LLM agent acts directly as the policy on a bimanual YAM arm setup: it reads camera observations through a tool interface, issues Cartesian and joint commands, and judges its own completion. Ten manipulation tasks (block stacking, die flipping, towel folding, part insertion and assembly, throwing), six models, three reasoning-effort… See the full description on the dataset page: https://huggingface.co/datasets/Agent-as-Policy/agent-as-policy.imageroboticsn<1K3 likes1.9k downloads24d agoHugging Face13PHY041 /gui-agent-outputimage1K<n<10K0 likes1.5k downloads6mo agoHugging Face14agentsea /wave-uiLICENSE image10K<n<100K27 likes1.3k downloads2y agoHugging Face15yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.3k downloads7mo agoHugging Face16ChrisDing1105 /unified-agent-trajectories Unified Benchmark Agent Trajectories Dataset release: v2.1.1 (2026-09-18)Record format: unified-agent-sft-v1 A growing collection of benchmark agent execution trajectories converted into one transparent, multimodal, tool-aware representation. These are complete recorded benchmark runs—not ordinary chat transcripts—including benchmark tasks, model reasoning and answers, tool calls, tool observations, runtime status, and benchmark scores when available. The directory layout is… See the full description on the dataset page: https://huggingface.co/datasets/ChrisDing1105/unified-agent-trajectories.imagetext-generation1K<n<10K3 likes987 downloads22d agoHugging Face17Agentic-MME /Agentic-MME Agentic-MME Dataset This is the official dataset for the Agentic-MME benchmark, featured in Hugging Face Daily Papers. Agentic-MME is a comprehensive benchmark designed to evaluate the abilities of multimodal agents in tool-use, web searching, and multi-step reasoning through visual clues. Usage You can load the dataset using the Hugging Face datasets library: from datasets import load_dataset dataset = load_dataset("Crystal1047/Agentic-MME") # To see the first record… See the full description on the dataset page: https://huggingface.co/datasets/Agentic-MME/Agentic-MME.imagevisual-question-answeringn<1K4 likes983 downloads6mo agoHugging Face18ppak10 /Agentic-SLS-ASTM Agentic-SLS-ASTM ASTM mechanical-test specimens (D638 tensile, D790 flex) printed on the Inova Mk1 SLS printer and pulled on an MTS / TestWorks Instron. Each row is a single specimen with full geometry, scalar results, stress–strain + raw DAQ curves, and — for SLS rows — FK references and an embedded snapshot of the upstream print profile from ppak10/Agentic-SLS-Database. Rows are self-contained for ML use: the full PrintProfile JSON is inlined, so features (material/energy… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-ASTM.documentn<1K0 likes937 downloads28d agoHugging Face19OpenDCAI /AgentFlow-DocDancer-benchmarkdocument10K<n<100K0 likes934 downloads7mo agoHugging Face20bubble65 /EMU-Agentic-PostTrain-Dataimage10K<n<100K4 likes921 downloads2mo agoHugging Face21mlfoundations-cua-dev /easyr1-agent-grounding-dataimage1M<n<10M1 likes878 downloads1y agoHugging Face22FLARE25-Agent-Xray /CBIS-DDSM-SEGimage1K<n<10K1 likes858 downloads1y agoHugging Face23orlando23 /failed_agent_trajectory Mobile Trajectory Verification Data This dataset is prepared for verifying how LoT agent can help us identify incorrect actions taken in mobile execution trajectory. As mentioned in our paper (see citation 1), We select 52 mobile agent trajectories with specific task goals from instruction- guided executions on MagicWand platform(see citation 2). Each folder has a execution trajectory that includes screenshots along this trajectories, json files describing the details of each… See the full description on the dataset page: https://huggingface.co/datasets/orlando23/failed_agent_trajectory.imagen<1K0 likes807 downloads2y agoHugging Face24ServiceNow /AgentHorizon AgentHorizon AgentHorizon is a benchmark for evaluating LLM judges of computer-use agents. Each item is a recorded trajectory of a GUI agent attempting a long-horizon desktop task (often 100 to 300+ steps), paired with a ground-truth label that says whether the agent actually completed the task. A judge reads the trajectory (instruction, action sequence, and screenshots) and predicts success or failure. Repository contents Path What it is… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/AgentHorizon.image100K<n<1M1 likes805 downloads1d agoHugging Face25ziggylott /agent-tts-libraryaudion<1K0 likes793 downloads2mo agoHugging Face26lobinni /apex-agents APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12 law)… See the full description on the dataset page: https://huggingface.co/datasets/lobinni/apex-agents.documentn<1K0 likes783 downloads8mo agoHugging Face27yuanyyaa /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/yuanyyaa/agent-reward-bench.imagerobotics1K<n<10K0 likes775 downloads6mo agoHugging Face28idleengine /apex-agents APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/apex-agents.documentn<1K0 likes763 downloads2mo agoHugging Face29agentsea /wave-ui-25k WaveUI-25k This dataset contains 25k examples of labeled UI elements. It is a subset of a collection of ~80k preprocessed examples assembled from the following sources: WebUI RoboFlow GroundUI-18K These datasets were preprocessed to have matching schemas and to filter out unwanted examples, such as duplicated, overlapping and low-quality datapoints. We also filtered out many text elements which were not in the main scope of this work. The WaveUI-25k dataset includes the original… See the full description on the dataset page: https://huggingface.co/datasets/agentsea/wave-ui-25k.image10K<n<100K40 likes741 downloads2y agoHugging Face30jzshared /agent_paper_reviewimage1K<n<10K0 likes701 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.