Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01semianalysisai /cc-traces-weka-062126 semianalysisai/cc-traces-weka-062126 WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:48:24 UTC via utils/agentic/build_weka_hf_dataset.py. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent sub-agent groups ≤ 10 Non-image rows only (image content excluded at source) Classifier calls excluded (max_tokens<=64 AND no tools → SUGGESTION MODE, title-gen… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126.texttext-generationn<1K13 likes45k downloads4mo agoHugging Face02semianalysisai /cc-traces-weka-062126-256k semianalysisai/cc-traces-weka-062126-256k WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:49:45 UTC via utils/agentic/build_weka_hf_dataset.py. Derived from semianalysisai/cc-traces-weka-062126 by applying the 256k per-request cap and preserving the surviving requests' relative timestamps. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126-256k.texttext-generationn<1K6 likes26k downloads4mo agoHugging Face03Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K19 likes6.7k downloads3d agoHugging Face04Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.8k downloads16d agoHugging Face05choucsan /mimo-claude-code-traces-1k MIMO Claude Code Traces MIMO Claude Code Traces is a collection of coding-agent trajectories in a Claude Code-style environment. Each record contains a user coding task, the full multi-turn message trace, available tool schemas, assistant reasoning fields, tool calls, tool outputs, and metadata such as model name, category, duration, cost, token usage, and whether the trace used tools. The traces were generated with mimo-v2.5-pro, MiMo's most capable model at the time of… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/mimo-claude-code-traces-1k.tabulartext-generation1K<n<10K11 likes2k downloads2mo agoHugging Face06Infatoshi /kernelbench-cuda-tracestabularn<1K2 likes1.9k downloads3d agoHugging Face07PRATHAM4567 /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/PRATHAM4567/Fable-5-traces.tabulartext-generation1K<n<10K0 likes1.4k downloads3mo agoHugging Face08uynitsuj /paper-sim-n128-traces WARP-RM paper simulation n=128 traces Canonical paper-A rollout traces for the two released simulation policies over the fixed 128-episode evaluation suite. They are the original trace pair used by the public evaluator self-test, rather than a later re-run. The archive contains 256 qpos traces: 128 vanilla-policy traces and 128 WARP-RM-policy traces. It is deliberately limited to evaluation output (no model weights, DINO features, or raw source-capture data). With the released… See the full description on the dataset page: https://huggingface.co/datasets/uynitsuj/paper-sim-n128-traces.textn<1K0 likes1.2k downloads3mo agoHugging Face09melissapan /swe-bench-lite-agent-traces-v14 AgentBRANE SWE-bench Lite Agent Traces v14 This release contains the 1,890 harness-native agent traces selected by the sealed SWE-bench Lite v14 publication record (1,379/1,890 resolved, 73.0%). It includes Claude Code, Codex, and Pi sessions across seven models and three replicates. No internal research notes are included. Load the observation table: from datasets import load_dataset traces = load_dataset("melissapan/swe-bench-lite-agent-traces-v14", split="train") Each row… See the full description on the dataset page: https://huggingface.co/datasets/melissapan/swe-bench-lite-agent-traces-v14.tabulartext-generation1K<n<10K0 likes997 downloads25d agoHugging Face10dchasap /spec_cpu_branch_tracestext10B<n<100B0 likes942 downloads2y agoHugging Face11shijunhao /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.tabulartext-generation1K<n<10K1 likes900 downloads4mo agoHugging Face12mlx-community /optiq-code-traces OptiQ Code Traces Gold-verified agentic software-engineering trajectories, produced by OptiQ Code, the terminal coding agent for local models on a Mac. Each trajectory is a full tool-calling run against a real repository bug, and every resolved label is set by executing the gold tests (FAIL_TO_PASS + PASS_TO_PASS) after applying the model's patch, never by the agent's own self-report. The dataset is 1,789 agent sessions in HuggingFace Session-Traces format (the agent-traces… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-code-traces.tabulartext-generation1K<n<10K4 likes881 downloads25d agoHugging Face13AlinCiocan /fable-5-claude-code-traces Fable 5 Claude Code Traces A full, scrubbed release of Fable 5 Claude Code session traces for researchers studying real coding-agent behavior: multi-turn prompts, assistant responses, tool calls, command output, retries, and session-level workflow metadata. This release keeps the full package intact: 18 sessions, 9,497 JSONL events, 0 excluded rows, and 0 quarantine files. The traces are preserved in the native agent-event format so they can be inspected in Hugging Face Agent… See the full description on the dataset page: https://huggingface.co/datasets/AlinCiocan/fable-5-claude-code-traces.tabulartext-generationn<1K7 likes842 downloads4mo agoHugging Face14sornnakub /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/sornnakub/Fable-5-traces.tabulartext-generation1K<n<10K1 likes816 downloads4mo agoHugging Face15Inferact /codex_swebenchpro_tracesThis is a dataset generated by real swebenchpro agentic workload trace + codex agent. 1. Eval Result Summary Metric Value Total trials 731 Successful trials 610 Failed trials 120 No data (skipped) 1 Passed 329 Pass rate (of successful) 53.9% Per-Repo Breakdown Repo Total Success Failed Passed Pass% ansible/ansible 96 93 3 60 65% internetarchive/openli 91 88 3 52 59% flipt-io/flipt85 82 3 26 32% qutebrowser/qutebrowse 79 78 1… See the full description on the dataset page: https://huggingface.co/datasets/Inferact/codex_swebenchpro_traces.textn<1K29 likes763 downloads5mo agoHugging Face16kira /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/kira/Fable-5-traces.tabulartext-generation1K<n<10K0 likes709 downloads4mo agoHugging Face17Flownium /zeus-30k-trace-dataset Zeus Qwen execution dataset Exactly 30,000 unique solver-input traces generated through the user-provided Qwen3.8-27B API. The endpoint reports this model ID; weights were not independently inspected. Thirty concurrent API workers were requested, with fallback to twenty after repeated overload errors. Categories: {"clarification":2000,"coding":6000,"documents":5000,"failure_recovery":2000,"maths":6000,"tool_use":4000,"writing":5000}. Contents traces:… See the full description on the dataset page: https://huggingface.co/datasets/Flownium/zeus-30k-trace-dataset.text100K<n<1M0 likes689 downloads14h agoHugging Face18evalstate /test-traces Test Traces Codex-style rollout JSONL traces exported from fast-agent for validation against the Hugging Face Agent Trace Viewer. tabularn<1K2 likes677 downloads5mo agoHugging Face19ApertureQA /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/ApertureQA/Fable-5-traces.tabulartext-generation1K<n<10K0 likes663 downloads3mo agoHugging Face20cfahlgren1 /Fable-5-tracesA simple dataset of the raw Fable 5 Claude session logs we could get our hands on before it was taken away (no clue if it's coming back). The raw trace files live in sessions/*.jsonl. Cache files, paste-cache files, shell history, and merged COT training exports are intentionally omitted so Hugging Face Datasets can load the repo through the agent-traces path. A pretty viewer for dataset:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Fable-5-traces.tabularn<1K19 likes613 downloads4mo agoHugging Face21armand0e /minimax-m3-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Minimax M3 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by minimax/minimax-m3. JSONL files: 31 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m3-claude-code-traces.tabulartext-generationn<1K13 likes603 downloads4mo agoHugging Face22armand0e /qwen3.7-max-pi-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Qwen3.7 Max Pi Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by qwen/qwen3.7-max. JSONL files: 47 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README. Use it… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-pi-traces.tabulartext-generationn<1K96 likes601 downloads5mo agoHugging Face23armand0e /kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 36 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.tabularn<1K4 likes595 downloads4mo agoHugging Face24PrimeIntellect /trace-cheating-recall-500 Trace Cheating Recall 500 This dataset contains 500 SWE-agent traces selected to evaluate whether an LLM judge detects observable solution leakage. It is the public data source for the trace-cheating-recall-500 Prime environment. The examples were derived from PrimeIntellect/int4-syn-gen-swe-glm53-bash-2026-09-02 at revision 0e7a9ecddce8de9ea8f8c369b2dd39411d6dee7a. Composition 500 unique traces, all labeled CHEATING 250 internet-retrieval cases 250 Git-history… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/trace-cheating-recall-500.texttext-classificationn<1K0 likes580 downloads1mo agoHugging Face25edbuildingstuff /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/edbuildingstuff/Fable-5-traces.tabulartext-generation1K<n<10K0 likes577 downloads3mo agoHugging Face26clem /hf-coding-tools-traces_april26 HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 32 sessions, one per (tool, model, effort, thinking) configuration 9,130 query → response turns total (≈18,260 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/clem/hf-coding-tools-traces_april26.tabularn<1K0 likes569 downloads5mo agoHugging Face27omlab /trace-bench TRACE Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding Timestamped video QA and proactive-response annotations for evaluating what a model knows, when it knows it, and how it responds. How TRACE works TRACE separates causal video delivery, model interaction, and scoring. The same public contract makes QA and Proactive Response results auditable across models: QA: answer a question using only the video evidence… See the full description on the dataset page: https://huggingface.co/datasets/omlab/trace-bench.tabularvisual-question-answering1K<n<10K2 likes547 downloads12d agoHugging Face28ILoveBuns /python-mental-execution-traces Python Mental Execution Traces A 12,000-row prompt/completion dataset for evaluating and training language models to mentally execute self-contained Python 3 snippets without running them. Completions provide the expected standard output together with a concise variable trace or explanation. Dataset structure The JSONL file contains two text fields: prompt: a Python mental-execution problem. completion: the expected stdout and concise reasoning or variable trace.… See the full description on the dataset page: https://huggingface.co/datasets/ILoveBuns/python-mental-execution-traces.texttext-generation10K<n<100K0 likes537 downloads2mo agoHugging Face29davidkling /hf-coding-tools-traces-run-april12 HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 8,875 query → response turns total (≈17,750 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-run-april12.tabularn<1K0 likes532 downloads5mo agoHugging Face30dacorvo /hf-hub-session-pi-traces dacorvo/hf-hub-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/hf-hub-session-captures. Both belong to the hf-hub-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes510 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.