datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lmcache-agentic-traces
LMCache Agentic Dataset Collection
A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache.
Motivation
Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.Nemotron-RL-Agentic-SWE-Pivot-v1
Dataset Description:
The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The dataset is a refactored version of the SWE-Gym and R2E-Gym datasets to support the NeMo Gym input format.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.freeinference_agentic_trace
FreeInference Agentic Trace
This dataset contains 16 weeks (2026-05-17 to 2026-09-05) of coding and assistant agents interacting with LLMs through the FreeInference gateway: 12,002 agent sessions from 267 accounts and 14 agent harnesses, with 1,186,582 LLM requests and 1,213,347 tool calls.
The dataset accompanies the paper From Requests to Sessions: A Large-Scale Characterization of Human-Driven Agentic Workloads. The code that reads it, replays its prefix cache and reproduces… See the full description on the dataset page: https://huggingface.co/datasets/harvardMadsys/freeinference_agentic_trace.H2EPR-Bench
H²EPR-Bench
An Evidence-Traceable Benchmark for Event-Process Reconstruction
H²EPR-Bench asks a demanding question: can a model reconstruct how a complex
real-world event unfolded, rather than merely summarize what happened? Given an
event specification and fixed multi-source evidence, a system produces a
hierarchical heterogeneous Event-Process Graph (EPG) that makes stages,
episodes, participants, actions, outcomes, relations, and evidence support
explicit.… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/H2EPR-Bench.Agentic-SLS-Telemetry
Inova-Mk1-Telemetry
Time-aligned printer-state recordings from Inova Mk1 SLS 3D print runs. One row per 10 Hz tick — the recorder's /state/snapshot poll — with the full sensor state snapshot (~64 columns: temperatures, position, power, lights) on every row, the nearest camera frame embedded inline when one fell within the prior 100 ms window, and any 1 kHz position-stream samples from that window collected as a nested list.
25 parquet files across builds spanning 2026-05 through… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-Telemetry.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.agentic-prompt-injection-5k
Agentic Prompt-Injection 5K
5,000 examples of agentic and indirect prompt injection. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. It focuses on the harder, agentic surface: tool/function abuse, RAG-document-embedded (indirect) injection, memory and trust-boundary poisoning, and approval/authority escalation.
Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec and OWASP AI Exchange / GenAI… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/agentic-prompt-injection-5k.Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖
Self-Taught Agentic Long Context Understanding (Arxiv).
AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass.
Installation Requirements
This codebase is largely based on OpenRLHF and Helmet, kudos to them.
The requirements are the same
pip install openrlhf
pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.agentic-polymarket
agentic-polymarket
38,915 settled Polymarket binary event markets with full hourly price curves, question text, resolution
terms, and ground truth outcomes. Prepared for research on "getting LLM agents to trade on prediction markets."
Companion code (backtest env + agent trading interface): see RSI-economy/shadow-market.
What this dataset solves
Historical price series cannot be used directly as backtest targets —— a recording does not react to
agent behavior: any… See the full description on the dataset page: https://huggingface.co/datasets/tennant/agentic-polymarket.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.public-agentic-cowork-benchmarks
Public Agentic Cowork Benchmarks
Download and restore large payloads
Large payloads are stored under archive-parts/ as 4 MiB LFS blocks to avoid unreliable large multipart transfers. PAYLOADS.json records every block hash and the exact original payload hash. Download this dataset repository with LFS payloads, then run:
python3 assemble_payloads.py
The script verifies each block and verifies that the reassembled archive or data file is byte-for-byte identical to… See the full description on the dataset page: https://huggingface.co/datasets/MitakaKuma/public-agentic-cowork-benchmarks.tb21-agentic-top10-full
Terminal-Bench 2.1: ten nearest training tasks per task, full pool
For each of the 89 Terminal-Bench 2.1 tasks, the ten training tasks that would most help a model solve it, chosen from 1,962,393 candidate tasks drawn from 32 public training corpora and benchmarks plus one private set. The method is the agentic selection from TB2.1 data-selection v2, run over a pool 25 times the size of v2's 76,786 tasks.
Files
Path
Contents
data/top10.parquet
890 rows… See the full description on the dataset page: https://huggingface.co/datasets/andylizf/tb21-agentic-top10-full.ViZDoom-Agentic-Rollouts
ViZDoom Agentic Rollouts
This dataset contains model-generated evaluation results for the pufanyi/ViZDoom benchmark. The benchmark repository contains only test specifications; this repository contains scores and rollout artifacts.
Both configs contain all 120 evaluated episodes: 10 seeds for each of the 12 default single-player ViZDoom environments. Browse the synchronized videos in the ViZDoom Agentic Demo.
Configs
qwen3.6-27b-5tics
The… See the full description on the dataset page: https://huggingface.co/datasets/pufanyi/ViZDoom-Agentic-Rollouts.GitHub-Agentic-PR-Dataset
GitHub Agentic PR Dataset
A large-scale dataset of ~2 million GitHub Pull Requests authored by AI coding agents (Claude Code, Cursor, GitHub Copilot, Devin) and human developers, complete with commits, file-level diffs, patches, bug-fix classification, review-depth metrics, and test-quality measures.
The GitHub Agentic PR Dataset is a research-grade corpus for studying how AI coding agents contribute to real-world open-source software, and how their pull requests compare to… See the full description on the dataset page: https://huggingface.co/datasets/mabujadallah/GitHub-Agentic-PR-Dataset.guidellm-agentic-coding-trajectories
GuideLLM agentic coding trajectories
A sampled serving-load benchmark derived from Thoughtworks agentic-coding-trajectories, for GuideLLM and an OpenAI-compatible /v1/chat/completions endpoint. There are 630 rows representing 481 unique source sessions, across the same 8turn, 24turn, and 48turn configurations as the earlier version.
The configuration names now refer to original logical steps, not always HTTP request counts. Native tool steps expand into a tool-call request and a… See the full description on the dataset page: https://huggingface.co/datasets/zetomatoz/guidellm-agentic-coding-trajectories.agentic-coding-trajectories
agentic-coding-trajectories
A unified, tokenized corpus of 15,000 multi-turn agentic-coding sessions (618K turns, 41 turns/session avg) drawn from three publicly-released upstream datasets. Built for benchmarking LLM serving systems on realistic multi-turn coding-agent workloads.
Why this exists
Most LLM serving benchmarks use single-shot prompts. Real coding agents work in long multi-turn loops where each turn appends to a growing prompt. This corpus captures that shape… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/agentic-coding-trajectories.agentic_vbench_rollouts
agentic-vbench calibration rollout archive
Complete, immutable copies of the calibration trajectories referenced from the
agentic-vbench task PRs. Unlike the copies committed in each task's
calibration/rollouts/, nothing here is elided: every sampled-frame payload the agent
saw is present.
Redaction policy (mechanical, applied identically to every file):
Absolute host filesystem paths from the runner machine are replaced with /workspace.
Capture-hardware and source-collection… See the full description on the dataset page: https://huggingface.co/datasets/yalesunxiatao/agentic_vbench_rollouts.frontier-agentic-corpus-v10.3
Frontier Agentic Corpus v10.3
A multi-teacher knowledge-distillation corpus of genuine multi-step agentic
trajectories, coding, mathematical, scientific, and data-analysis reasoning,
built exclusively from currently frontier-eligible teacher models under a
strict, live-verified, two-leaderboard gate — released with row-level
provenance and a separate public/redistribution split.
Metric
Value
Total tokens
1,095,987,979 (chars/4 estimate)
Total examples
430,532… See the full description on the dataset page: https://huggingface.co/datasets/superisaac1262/frontier-agentic-corpus-v10.3.agentic_review
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Razvan27/agentic_review.PortBench-Market
PortBench Market Base Dataset
Dataset Description
A ten-year (Jan 2015–Dec 2025) daily financial dataset covering 183 instruments across six heterogeneous asset classes, designed for multi-asset portfolio management research and LLM evaluation.
Asset Coverage
Asset Class
Instruments
Data Fields
Sources
Equities
126
OHLCV + return
Yahoo Finance (ETFs: broad market, sector, factor, international)
Bonds
16
Close + return (ETFs); yield… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-Market.matrix-game-evaldeepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-0813-agentic.agentic-dataset-v2
Agentic Dataset v6 Dynamic
Dynamically allocated SFT dataset.
Requested target:
100,000
Actual target:
100,000
Domain distribution
coding: 32,200
agent: 15,235
tool: 22,565
math: 14,000
instruction: 10,000
stem: 6,000
Reliability
anti_hallucination:
30,512
tool_call_rows:
22,265
tool_result_rows:
21,972
tool_grounded_rows:
21,971
verified_rows:
18,384
uncertainty_rows:
1,058
fake_tool_claims:
0
Length
P50 chars:
2,038
P95 chars:
14… See the full description on the dataset page: https://huggingface.co/datasets/usernamebetter/agentic-dataset-v2.PyFi-600K
Dataset Card for PyFi-600K
This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents.
AgenticFinLab/PyFi-600K/
├── README.md # Dataset documentation and description
├── images.zip # Compressed image files
├── PyFi-600K-dataset.csv # Q&A pairs in CSV format
├── PyFi-600K-dataset.json # Q&A pairs in JSON format
├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset
└──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.slimder-qwen38-agentic-ream-runs-20260901
