datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tau2-Bench-Airline-With-Code-Agents
Dataset Card for a Code Agent Version of Tau Bench 2 Airline
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant. The dataset is based on the Airline environment from Tau^2 Bench and contains traces from both the original version and a version made at Snorkel AI using code agents to solve the same tasks (indicator in the version field; details below).
Curated by: Snorkel AI… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Airline-With-Code-Agents.Tau2-Bench-Verified-Airline-With-Code-Agents
Dataset Card for a Code Agent Version of Tau Bench 2 Airline
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant, along with the original verion of the tasks with more bespoke tools.
The dataset is based on a verified version of the Airline environment from Sierra.ai's Tau^2 Bench with the verified version from Amazon AGI group here.
You can find an earlier version of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Verified-Airline-With-Code-Agents.codeagent-tracesai-code-generation-swe-agents-2026
💻 AI Code Generation, SWE Agents & Program Synthesis Dataset (2026 Edition)
A structured research dataset featuring 3,181 domain-verified research papers and 771 official code repositories focused on Autonomous Software Engineering Agents (SWE-bench), Program Synthesis, DeepSeek-Coder-V2, Qwen2.5-Coder, Test-Driven Code Repair, Self-Healing Software, AST Semantic Modeling, and Formal Logic Verification (2023–2026).
Built with Universal Scientific Engine V17.1 Gold, providing 47… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/ai-code-generation-swe-agents-2026.agent-traces-code-review-pipeline
Agent Traces: code-review-pipeline
Synthetic multi-agent workflow traces with LLM-enriched content for the code-review-pipeline domain.
Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns.
What is this dataset?
This dataset contains 2,035 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes:
Agent reasoning — chain-of-thought for each agent step
LLM… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/agent-traces-code-review-pipeline.misc-merged-claude-code-traces-v1
MISC Unification of Public Claude Code Traces
A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format.
Dataset Description
This dataset contains real Claude API interaction traces capturing software engineering workflows including:
Code generation and modification
Bug fixing and… See the full description on the dataset page: https://huggingface.co/datasets/agent-data/misc-merged-claude-code-traces-v1.hermes-function-calling-v1-formatted-code-agenthermes-codeagentAFM-CodeAgent-RL-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-CodeAgent-RL-Dataset.agentic-code-search-rollouts-sample10codeagent-tracesarxiv-to-code-agentic-tool-calling
arxiv-to-code-agentic-tool-calling
Multi-turn tool-calling dataset where an assistant implements ML papers in PyTorch through file-creation and command-execution tool calls.
Built from lucidrains' (Phil Wang) open-source paper implementations. There are ~217 repositories on Codeberg, each implementing a different ML paper. This dataset reverse-engineers those into synthetic coding conversations.
What's in it
199 conversations, each covering one repository. Every… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/arxiv-to-code-agentic-tool-calling.codeagent-traces-answerscode_agent_datacode-contests-sandboxes-traces-terminus-2EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPOcodeagent-traces-tool-rolecodeagent-traces-user-rolemem_agent-model_based-memagent-1-5b-separate-step720-infbench-code-debug-test-c27000-t4096-10s-agenerated_code-Agentmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-1000s-a-fullcontextmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-1000s-a-nocontextmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c31000-t4096-1000s-agnosticmem_agent-model_based-rl-memoryagent-14b-infbench-code-debug-test-c31000-t4096-1000s-agnosticmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c27000-t4096-1000s-agnosticmem_agent-model_based-rl-memoryagent-14b-infbench-code-debug-test-c27000-t4096-1000s-agnosticmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-1000s-agnosticmem_agent-model_based-memagent-1-5b-step1024-infbench-code-debug-test-c8192-t4096-10s-agnosticmem_agent-model_based-qwen3-1-5b-oldgrpo-2086-infbench-code-debug-test-c8192-t4096-1000s-agnostimem_agent-model_based-memagent-1-5b-separate-step720-infbench-code-debug-test-c27000-t4096-1000s
