Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01semianalysisai /cc-traces-weka-062126 semianalysisai/cc-traces-weka-062126 WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:48:24 UTC via utils/agentic/build_weka_hf_dataset.py. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent sub-agent groups ≤ 10 Non-image rows only (image content excluded at source) Classifier calls excluded (max_tokens<=64 AND no tools → SUGGESTION MODE, title-gen… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126.texttext-generationn<1K12 likes41k downloads4mo agoHugging Face02nvidia /Open-SWE-Traces Open-SWE-Traces: Advancing Distillation for Software Engineering Agents 🚨 What's New [09/26] Release v1.2: Added new agent trajectories generated by Qwen3.8-27B for mini-swe-agent. Trajectories for OpenCode and Claude Code harnesses will be released soon. [08/26] Release v1.1: Added new agent trajectories generated by DeepSeek-V4-Flash and Qwen3.6-27B across OpenHands, SWE-agent, and mini-swe-agent harnesses. [06/21] Release v1.0: Released 207k agent… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Open-SWE-Traces.text100K<n<1M140 likes26k downloads14d agoHugging Face03semianalysisai /cc-traces-weka-062126-256k semianalysisai/cc-traces-weka-062126-256k WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:49:45 UTC via utils/agentic/build_weka_hf_dataset.py. Derived from semianalysisai/cc-traces-weka-062126 by applying the 256k per-request cap and preserving the surviving requests' relative timestamps. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126-256k.texttext-generationn<1K6 likes26k downloads4mo agoHugging Face04yinhuankuang /rl-game-traces-rise-of-the-tomb-raider 古墓丽影:崛起 This public dataset repository contains gameplay trace data uploaded from F:\古墓丽影:崛起. Contents Files: 647 Total local size: 496.11 GB Generated: 2026-06-06T01:03:39+00:00 File Types .jsonl: 196 .json: 149 .png: 129 .parquet: 49 .mkv: 49 .txt: 49 .jpg: 15 .exe: 11 Notes This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs. The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-rise-of-the-tomb-raider.imagereinforcement-learning0 likes23k downloads4mo agoHugging Face05cot-leaderboard /cot-eval-traces-2.0text1M<n<10M9 likes15k downloads2y agoHugging Face06flashinfer-ai /flashinfer-trace FlashInfer Trace We provide an official dataset called FlashInfer Trace with kernels and workloads in real-world AI system deployment environments. FlashInfer-Bench can use this dataset to measure and compare the performance of kernels. It follows the FlashInfer Trace Schema. It is organized as follows: flashinfer_trace/ # Here ├── definitions/ └── workloads/ flashinfer-trace/ # On Hugging Face ├── solutions/ └── traces/ Example solutions and traces directories, featuring… See the full description on the dataset page: https://huggingface.co/datasets/flashinfer-ai/flashinfer-trace.20 likes10k downloads4mo agoHugging Face07imo2026-challenge /chankhavu-imo-reasoning-traces0 likes8.2k downloads2mo agoHugging Face08dacorvo /funes-nvidia-Open-SWE-Traces Funes recall store — NVIDIA Open-SWE-Traces (resolved) A funes recall store built by indexing the resolved==1 trajectories of nvidia/Open-SWE-Traces (65244 sessions, across both harnesses — SWE-agent and OpenHands — and both models, Minimax-M2.5 and Qwen3.5-122B). What this is This is not a raw trace dataset — it is a pre-built funes index: the source trajectories chunked into content blocks and embedded, stored as a Lance table (chunks.lance). Source… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-nvidia-Open-SWE-Traces.10M<n<100M0 likes7.8k downloads5d agoHugging Face09Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K18 likes6.6k downloads7d agoHugging Face10Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes6.2k downloads12d agoHugging Face11masterpieceexternal /gpt-oss-20b-moe-expert-power-traces-320k GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer) This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100. What is recorded Each trace corresponds to one capture trial where: A fixed expert id is selected (expert_00 ... expert_31). A random hidden-state tensor is generated once per trial. The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.audio-classification100K<n<1M0 likes5.1k downloads4mo agoHugging Face12VCG-EAI /TraceSpatial-Trace TraceSpatial-Trace Project · Paper · Code TraceSpatial-Trace is the RGB-and-QA tracing subset of TraceSpatial, introduced by the RoboTracer project. It supports learning to translate language instructions into spatial waypoints for object manipulation and robot end-effector motion. This release contains 517,215 referenced RGB images and 3,623,880 question–answer pairs, extracted from 1,067,822 original conversation records across AgiBot, CA-1M, DROID, RoboTwin, and ScanNet. It… See the full description on the dataset page: https://huggingface.co/datasets/VCG-EAI/TraceSpatial-Trace.imagevisual-question-answering1M<n<10M4 likes4.9k downloads2d agoHugging Face13jonathanyin /aime_1983_2023_deepseek-r1_traces_16384tabularn<1K0 likes4.8k downloads1y agoHugging Face14osazuwa /wham3d-cts-trace-markov-rollouts-20260923 3D WHAM Trace+Markov cached video continuations Growing, publicly downloadable video-backed rollout dataset for the trace_markov ablation arm at training step 500, seed 20260921. Each selected frozen observational prefix has exactly one cached continuation in each of two routes: null-intervention observational and diagnostic do(X=1). Prefix IDs are sorted once within each (W,Z) stratum and remain balanced as the cohort grows. The direct route of an observationally trained arm is… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/wham3d-cts-trace-markov-rollouts-20260923.video1K<n<10K0 likes3.8k downloads10d agoHugging Face15osazuwa /wham3d-cts-trace-rollouts-20260923 3D WHAM Trace cached video continuations Growing, publicly downloadable video-backed rollout dataset for the trace ablation arm at training step 500, seed 20260921. Each selected frozen observational prefix has exactly one cached continuation in each of two routes: null-intervention observational and diagnostic do(X=1). Prefix IDs are sorted once within each (W,Z) stratum and remain balanced as the cohort grows. The direct route of an observationally trained arm is a diagnostic;… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/wham3d-cts-trace-rollouts-20260923.video1K<n<10K0 likes3.7k downloads10d agoHugging Face16nlile /misc-merged-claude-code-traces-v1 MISC Unification of Public Claude Code Traces A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format. Dataset Description This dataset contains real Claude API interaction traces capturing software engineering workflows including: Code generation and modification Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.text10K<n<100K19 likes3.7k downloads10mo agoHugging Face17devichand /schedulerlens-traces1 likes3.5k downloads3mo agoHugging Face18dungnv /qwen36-27b-length-traces Qwen3.6-27B generation-length prediction: heads, calibrations and workloads Artifacts for conformal length-aware LLM scheduling on Qwen/Qwen3.6-27B — predicting a request's remaining generation length from a hidden layer during decoding, wrapping it in a split-conformal interval, and scheduling with SRPT inside vLLM. Extends TRAIL (Don't Stop Me Now, ICLR'25) to a hybrid-attention reasoning model. This repo contains the derived artifacts, not the raw activations. The 3250… See the full description on the dataset page: https://huggingface.co/datasets/dungnv/qwen36-27b-length-traces.text-generation0 likes3k downloads2mo agoHugging Face19Ryze242005 /vllm-traces-v2tabularn<1K0 likes2.8k downloads12d agoHugging Face20FrontisAI /NatureBench-traces NatureBench-traces NatureBench-traces contains the full solving process of coding agents on the 90 tasks of NatureBench. The task packages themselves (task brief, data, evaluator, SOTA scores) live in the sibling repository FrontisAI/NatureBench. Harbor-compatible task packages are available in FrontisAI/NatureBench-Harbor. The traces released here were collected with NatureBench's native task format, not from the Harbor tasks. This repository releases only the process traces:… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/NatureBench-traces.tabular1K<n<10K3 likes2.6k downloads1mo agoHugging Face21experiential-labs /wmo-terminal-tasks-traces terminal-tasks — real agent-environment traces Computer-use agent runs in real terminal containers: bash commands and their true outputs from live task environments. Every trace is a REAL run: an LLM agent stepping against the actual benchmark environment, with each transition (tool call → true environment observation) recorded as OpenTelemetry GenAI spans (traces.otel.jsonl, one span per line). Captured by world-model-harness's environment-capture package, which also holds the… See the full description on the dataset page: https://huggingface.co/datasets/experiential-labs/wmo-terminal-tasks-traces.2 likes2.5k downloads2mo agoHugging Face22experiential-labs /wmo-bird-sql-traces bird-sql — real agent-environment traces Text-to-SQL over real SQLite databases: the agent explores a copy of the task's database and schema, then submits a SQL query. Every trace is a REAL run: an LLM agent stepping against the actual benchmark environment, with each transition (tool call → true environment observation) recorded as OpenTelemetry GenAI spans (traces.otel.jsonl, one span per line). Captured by world-model-harness's environment-capture package, which also holds… See the full description on the dataset page: https://huggingface.co/datasets/experiential-labs/wmo-bird-sql-traces.0 likes2.5k downloads2mo agoHugging Face23Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K1 likes2.4k downloads3mo agoHugging Face24tonychenxyz /sweeperbench-traces Sweeperbench run database runs.csv is the append-only raw table. Each row describes one prediction, core evaluation, or preservation evaluation. schema.json documents the fields. Primary key: run_id (UUID). Each top-level UUID directory is exactly that run, with matching metadata.json, artifact manifest, prompts, traces, verdicts and any collected recordings. Relationships: experiment_id groups a launch; evaluations reference the exact prediction_run_id. Model, harness… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/sweeperbench-traces.tabular10K<n<100K0 likes2.3k downloads10d agoHugging Face25TeichAI /Ox-Alpha-Pi-TracesThis dataset was generated using teich by TeichAI Ox-Alpha Pi Agent Coding Traces This directory contains raw agent trace files generated by teich. JSONL files: 2247 Model metadata: stealth/ox-alpha Domains and prompt distribution Topic Traces Games & simulation (headless) 196 Frontend & Node-testable web 159 Health & medicine informatics 139 ML & scientific computing (CPU) 123 Data analysis & reporting 122 Computational biology & chemistry… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-Pi-Traces.text-generation8 likes2.3k downloads1mo agoHugging Face26mlx-community /optiq-lab-traces OptiQ Lab Traces Research and tool-calling sessions produced by OptiQ Lab, the local web UI that ships with mlx-optiq. Each session is a complete run: a deep-research report built from live web sources, or a multi-turn agent loop driving the Lab's own sandboxed tools. The dataset is 866 sessions in HuggingFace Session-Traces format (the agent-traces viewer). Each .jsonl file is one session: a header line carrying the run's metadata, then one message per turn. The two… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-lab-traces.tabulartext-generationn<1K1 likes2.2k downloads2mo agoHugging Face27RESMP-DEV /Fable-GPT-5.5-Distillation-Traces Agent Traces Curated 2026 (v3 Merged) A unified distillation corpus of 9,057,143 records spanning agentic coding traces, math/code/science reasoning, tool-use trajectories, and preference data. 8,876,012 train + 181,131 eval, stratified by source. What this is This is the v3 merged corpus that supersedes both v1 and v2 of this dataset. It combines five major source groups through a unified normalization pipeline: Original v2 RESMP-DEV (de-fragmented, re-deduped):… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/Fable-GPT-5.5-Distillation-Traces.texttext-generation1M<n<10M10 likes2.2k downloads3mo agoHugging Face28core12345 /MoE_expert_selection_tracegated 📖 Introduction This repository serves as a supplement to our paper "Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference". It contains expert selection profiling traces of four top-tier MoE LLMs ranging from 235B to 1T (DeepSeek-R1, Kimi-K2-Thinking, Llama4-Marverick, and Qwen3-235B) across multiple benchmarks. For each query or request, we log the activated expert ID of every model layer of every generated token. We provide analyses and… See the full description on the dataset page: https://huggingface.co/datasets/core12345/MoE_expert_selection_trace.28 likes2.2k downloads5mo agoHugging Face29agent-evals /hal_traces8 likes2.2k downloads8mo agoHugging Face30MaxDevv /real-pi-coding-agent-traces-sessions Real Pi Coding Agent Traces Sessions An aggregated dataset of real human–AI coding agent sessions, collected from 21 independently published Hugging Face datasets and hand-filtered to exclude synthetic or AI-generated content. Every session is an unedited (but redacted) trace of a real person using pi — an open-source AI coding agent harness — to build, debug, and ship real open-source software. Real prompts, real tool calls, real errors, real backtracking. Why this… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/real-pi-coding-agent-traces-sessions.text-generation1K<n<10K5 likes2.2k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.