Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01semianalysisai /cc-traces-weka-062126 semianalysisai/cc-traces-weka-062126 WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:48:24 UTC via utils/agentic/build_weka_hf_dataset.py. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent sub-agent groups ≤ 10 Non-image rows only (image content excluded at source) Classifier calls excluded (max_tokens<=64 AND no tools → SUGGESTION MODE, title-gen… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126.texttext-generationn<1K13 likes45k downloads4mo agoHugging Face02semianalysisai /cc-traces-weka-062126-256k semianalysisai/cc-traces-weka-062126-256k WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:49:45 UTC via utils/agentic/build_weka_hf_dataset.py. Derived from semianalysisai/cc-traces-weka-062126 by applying the 256k per-request cap and preserving the surviving requests' relative timestamps. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126-256k.texttext-generationn<1K6 likes26k downloads4mo agoHugging Face03nvidia /Open-SWE-Traces Open-SWE-Traces: Advancing Distillation for Software Engineering Agents 🚨 What's New [09/26] Release v1.2: Added new agent trajectories generated by Qwen3.8-27B for mini-swe-agent. Trajectories for OpenCode and Claude Code harnesses will be released soon. [08/26] Release v1.1: Added new agent trajectories generated by DeepSeek-V4-Flash-0731 and Qwen3.6-27B across OpenHands, SWE-agent, and mini-swe-agent harnesses. [06/21] Release v1.0: Released 207k agent… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Open-SWE-Traces.text100K<n<1M146 likes25k downloads2d agoHugging Face04cot-leaderboard /cot-eval-traces-2.0text1M<n<10M9 likes13k downloads2y agoHugging Face05Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K19 likes6.7k downloads3d agoHugging Face06Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.8k downloads16d agoHugging Face07VCG-EAI /TraceSpatial-Trace TraceSpatial-Trace Project · Paper · Code TraceSpatial-Trace is the RGB-and-QA tracing subset of TraceSpatial, introduced by the RoboTracer project. It supports learning to translate language instructions into spatial waypoints for object manipulation and robot end-effector motion. This release contains 517,215 referenced RGB images and 3,623,880 question–answer pairs, extracted from 1,067,822 original conversation records across AgiBot, CA-1M, DROID, RoboTwin, and ScanNet. It… See the full description on the dataset page: https://huggingface.co/datasets/VCG-EAI/TraceSpatial-Trace.imagevisual-question-answering1M<n<10M4 likes5.3k downloads7d agoHugging Face08jonathanyin /aime_1983_2023_deepseek-r1_traces_16384tabularn<1K0 likes4.9k downloads1y agoHugging Face09zlab-princeton /SWEeper-Bench-traces 🧹 SWEeper-Bench agent traces Paper · Blog · Code · Benchmark This dataset has every run behind the SWEeper-Bench paper: 5,200 runs from 42 agent configurations. For each run you get the prompt the coding agent saw, its full trajectory, the shell commands it ran against the app, its patch, and the browser verifier's trajectories and verdicts. Main results All 15 agents ran on the 200 tasks with the open-ended prompt and the browser prompt template. A task… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/SWEeper-Bench-traces.tabular1K<n<10K1 likes4.3k downloads3d agoHugging Face10tonychenxyz /sweeperbench-traces Sweeperbench run database runs.csv is the append-only raw table. Each row describes one prediction, core evaluation, or preservation evaluation. schema.json documents the fields. Primary key: run_id (UUID). Each top-level UUID directory is exactly that run, with matching metadata.json, artifact manifest, prompts, traces, verdicts and any collected recordings. Relationships: experiment_id groups a launch; evaluations reference the exact prediction_run_id. Model, harness… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/sweeperbench-traces.tabular10K<n<100K0 likes4k downloads14d agoHugging Face11nlile /misc-merged-claude-code-traces-v1 MISC Unification of Public Claude Code Traces A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format. Dataset Description This dataset contains real Claude API interaction traces capturing software engineering workflows including: Code generation and modification Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.text10K<n<100K19 likes3.8k downloads10mo agoHugging Face12Ryze242005 /vllm-traces-v2tabularn<1K0 likes2.8k downloads16d agoHugging Face13FrontisAI /NatureBench-traces NatureBench-traces NatureBench-traces contains the full solving process of coding agents on the 90 tasks of NatureBench. The task packages themselves (task brief, data, evaluator, SOTA scores) live in the sibling repository FrontisAI/NatureBench. Harbor-compatible task packages are available in FrontisAI/NatureBench-Harbor. The traces released here were collected with NatureBench's native task format, not from the Harbor tasks. This repository releases only the process traces:… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/NatureBench-traces.tabular1K<n<10K3 likes2.7k downloads1mo agoHugging Face14Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K2 likes2.6k downloads3mo agoHugging Face15Mithilss /distillation-stream-traces Distillation stream traces Tool-call traces from open models working on agent benchmarks. The AI Alignment @ Illinois distillation stream records them to test whether a student model copies how its teacher acts with tools. Each run is one fresh attempt at one task. Runs Folder Model Tasks Solved Date qwen3.8-27b/first-traces-v0 Qwen/Qwen3.8-27B, revision 1d4bf0f2 10 from SWE-bench Verified 9 2026-10-01 qwen3.8-27b/four-benchmarks-v0 Qwen/Qwen3.8-27B 2… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/distillation-stream-traces.text0 likes2.4k downloads19h agoHugging Face16mlx-community /optiq-lab-traces OptiQ Lab Traces Research and tool-calling sessions produced by OptiQ Lab, the local web UI that ships with mlx-optiq. Each session is a complete run: a deep-research report built from live web sources, or a multi-turn agent loop driving the Lab's own sandboxed tools. The dataset is 866 sessions in HuggingFace Session-Traces format (the agent-traces viewer). Each .jsonl file is one session: a header line carrying the run's metadata, then one message per turn. The two… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-lab-traces.tabulartext-generationn<1K1 likes2.4k downloads2mo agoHugging Face17jonathanyin /aime_1983_2023_qwq-32b_traces_16384tabularn<1K0 likes2.2k downloads1y agoHugging Face18amanutej /trustworthy-biology-agents-traces Trustworthy Biology Agents — Run Traces Raw execution traces from 1,329 agent runs across three coding agents on three biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed trace bundle for the study in manu-tej/ai-scientists; the write-up lives in that repo's RESULTS.md. The motivating question is not only whether an agent reaches the right answer, but whether it behaves like a trustworthy analyst when the task is ambiguous, under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.tabular1K<n<10K0 likes2.2k downloads3mo agoHugging Face19RESMP-DEV /Fable-GPT-5.5-Distillation-Traces Agent Traces Curated 2026 (v3 Merged) A unified distillation corpus of 9,057,143 records spanning agentic coding traces, math/code/science reasoning, tool-use trajectories, and preference data. 8,876,012 train + 181,131 eval, stratified by source. What this is This is the v3 merged corpus that supersedes both v1 and v2 of this dataset. It combines five major source groups through a unified normalization pipeline: Original v2 RESMP-DEV (de-fragmented, re-deduped):… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/Fable-GPT-5.5-Distillation-Traces.texttext-generation1M<n<10M10 likes2.2k downloads3mo agoHugging Face20sammshen /lmcache-agentic-traces LMCache Agentic Dataset Collection A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache. Motivation Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.tabulartext-generation10K<n<100K17 likes2.1k downloads4mo agoHugging Face21choucsan /mimo-claude-code-traces-1k MIMO Claude Code Traces MIMO Claude Code Traces is a collection of coding-agent trajectories in a Claude Code-style environment. Each record contains a user coding task, the full multi-turn message trace, available tool schemas, assistant reasoning fields, tool calls, tool outputs, and metadata such as model name, category, duration, cost, token usage, and whether the trace used tools. The traces were generated with mimo-v2.5-pro, MiMo's most capable model at the time of… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/mimo-claude-code-traces-1k.tabulartext-generation1K<n<10K11 likes2k downloads2mo agoHugging Face22mlfoundations-dev /terminal-bench-traces-localtext1K<n<10K0 likes2k downloads1y agoHugging Face23Infatoshi /kernelbench-cuda-tracestabularn<1K2 likes1.9k downloads3d agoHugging Face24Lottolabs /terminal-bench-2.1-qwen3.8-27b-traces Terminal-Bench 2.1 traces: Qwen3.8-27B-GPTQ-4bit, xhigh / medium / low / off Complete agent trajectories, verifier output, timing and token usage for all 89 Terminal-Bench 2.1 tasks run locally with btbtyler09/Qwen3.8-27B-GPTQ-4bit on 2× RTX 3090, plus the adaptive fallback reruns at lower reasoning effort. Headline result: 62/89 (69.66%) at xhigh in a single clean pass. Cumulative best-of across xhigh → medium → low → off fallbacks: 70/89 (78.65%). The second number is not a… See the full description on the dataset page: https://huggingface.co/datasets/Lottolabs/terminal-bench-2.1-qwen3.8-27b-traces.text1K<n<10K1 likes1.6k downloads1mo agoHugging Face25trace-commons /agent-traces Trace Commons — Agent Traces Trace Commons is one open, public dataset of coding-agent sessions — the back-and-forth between a developer and an AI coding agent, including prompts, model responses, tool calls, and command output — contributed voluntarily as an open resource for studying, evaluating, and building on how these agents actually work. Every trace here was donated only from a public, open-source repository, was anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.tabulartext-generationn<1K36 likes1.5k downloads4mo agoHugging Face26DSFFGFG456 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K4 likes1.5k downloads2mo agoHugging Face27PRATHAM4567 /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/PRATHAM4567/Fable-5-traces.tabulartext-generation1K<n<10K0 likes1.4k downloads3mo agoHugging Face28greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.4k downloads2mo agoHugging Face29tjhunter /climate-tracetabular100M<n<1B2 likes1.4k downloads2y agoHugging Face30Crownelius /Complete-FABLE.5-traces-2M license: mit pretty_name: Claude Library — Fable 5 · Opus · Sonnet annotations_creators: machine-generated language: en language_creators: found machine-generated multilinguality: monolingual size_categories: 10K<n<100K task_categories: text-generation task_ids: language-modeling tags: agent-traces claude claude-fable-5 claude-opus claude-sonnet chain-of-thought tool-use coding-agents content-verified maintained-mirror deduplicated parquet configs: config_name:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Complete-FABLE.5-traces-2M.tabular100K<n<1M153 likes1.3k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.