datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepScaleR-Preview-Dataset
Data
Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from:
AIME (American Invitational Mathematics Examination) problems (1984-2023)
AMC (American Mathematics Competition) problems (prior to 2023)
Omni-MATH dataset
Still dataset
Format
Each row in the JSON dataset contains:
problem: The mathematical question text, formatted with LaTeX notation.
solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.Nemotron-SFT-Agentic-v2
Dataset Description
The Nemotron-SFT-Agentic-v2 dataset is a collection of synthetic single-turn and multi-turn tool-use trajectories designed to strengthen models’ capabilities as interactive, tool-using agents. It targets tasks where the model must decompose user goals, decide when to call tools, and reason over tool outputs to complete tasks reliably and safely.
This dataset is ready for commercial use.
The dataset consolidates three internally curated components (described… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Agentic-v2.DeepCoder-Preview-Dataset
Data
Our training dataset consists of 24K problems paired with their test cases:
7.5K TACO Verified problems.
16K verified coding problems from PrimeIntellect’s SYNTHETIC-1.
600 LiveCodeBench (v5) problems submitted between May 1, 2023 and July 31, 2024.
Our test dataset consists of:
LiveCodeBench (v5) problems between August 1, 2024 and February 1, 2025.
Codeforces problems from Qwen/CodeElo.
Format
Each row in the dataset contains:
problem: The coding problem… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepCoder-Preview-Dataset.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.formal-math-autoformalization
Formal Math Autoformalization Dataset
A growing, CC0 public-domain corpus of ⟨natural-language statement ↔ Lean 4 statement + proof⟩ pairs, contributed through the Agentic Commons network.
Why this is scarce data. Mathlib already contains millions of proven Lean theorems — but as bare Lean, with no paired natural language:
theorem add_comm (a b : ℕ) : a + b = b + a := ... -- no "addition on naturals is commutative" attached
The scarce, valuable artifact is the pairing of the… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/formal-math-autoformalization.Nemotron-Agentic-v1
Dataset Description:
The Nemotron-Agentic-Tool-Use-v1 dataset is designed to strengthen models’ capabilities as interactive, tool-using agents. It focuses on multi-turn conversations where language models decompose user goals, decide when to call tools, and reason over tool outputs to complete tasks reliably and safely.
This dataset is ready for commercial use.
The Nemotron-Agentic-Tool-Use-v1 dataset contains the following subsets:
Interactive Agent
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Agentic-v1.ai-research-berkeley-agentic-verification-harness-opt
Agentic Verification Meta-Verifier Traces
This public, manually gated Dataset repository stores immutable phase snapshots
from Meta-Verifier experiments.
Each run is organized as:
experiments/<theme>/<method>/<run>/phases/
train/ # Solver, delegated-verifier, Proposer/Reflector, harness population
val/ # Full validation traces, metrics, and selected frozen harness
heldout/ # Claimed held-out stage, all scheduled cells, and final scores
Access requests are… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-agentic-verification-harness-opt.AgenticOCR-SFT
AgenticOCR SFT Training Data
Supervised fine-tuning data for the AgenticOCR project.
The dataset contains 7,631 training records in sft_combined_0422.json. Image paths in each record are relative to the repository root and point into sft_images/.
agentic-env-index
Agentic Env Index
An index of materials for building agent environments: not a list of environments, and no content is copied here.
It covers 990,633,885 public files and records (4.04 PB) from 88 sources. Each row says where one item can be downloaded, how large it is, which hashes verify it and under which license it is offered. The files themselves stay at their sources.
Use
This repository gets new commits every day. Pin one commit so that the same query gives… See the full description on the dataset page: https://huggingface.co/datasets/Battam/agentic-env-index.AgenticDataBench
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Project Page | GitHub | Paper
AgenticDataBench is a comprehensive benchmark for evaluating LLM-based data agents that automate real-world data science workflows. It addresses the lack of rigorous evaluation by providing diverse, realistic tasks with fine-grained ground-truth labels.
The benchmark spans 15 domains, including real B2B fintech use cases, and is structured around reusable data science skills—core… See the full description on the dataset page: https://huggingface.co/datasets/shawnzzzh/AgenticDataBench.sol-max-opusnode-data
sol-max-opusnode-data
Training data built by the AgentPTB arm for cell sol-max-opusnode — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max-opusnode.h*, and the companion to the run record in agentic-ptb/sol-max-opusnode-record.
field
value
plot cell
sol-max-opusnode
driver
Codex / gpt-5.6-sol… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-opusnode-data.agentic-ai-options-resultsNemotron-RL-Agentic-Indirect-Prompt-Injection-v1
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1
Dataset Description:
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection (IPI) attacks hidden inside tool-returned environment data. In each record, the agent receives a benign user request that requires calling a read tool whose output contains an adversarial instruction disguised as legitimate domain content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.Nemotron-RL-Agentic-Function-Calling-Pivot-v1
Dataset Description:
This is a RL dataset for general function-calling by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1.Nemotron-RL-Agentic-Terminal-Pivot-v1
Dataset Description
The Nemotron-RL-Agentic-Terminal-Pivot-v1 dataset provides training samples for reinforcement learning of command-line ("terminal use") LLM agents with the terminus_judge environment in NeMo Gym.
Each record is a single agent decision point extracted from a successful agent trajectory on a terminal task:
responses_create_params.input — the prompt: the task instruction plus the terminal interaction history (prior agent actions and terminal outputs) up to the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1.lmcache-agentic-traces
LMCache Agentic Dataset Collection
A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache.
Motivation
Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.Nemotron-RL-Agentic-SWE-Pivot-v1
Dataset Description:
The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The dataset is a refactored version of the SWE-Gym and R2E-Gym datasets to support the NeMo Gym input format.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.project2-agentic-langdata-es
Agentic Language Learning — ES Dataset
Auto-prepared via Data Ingestion & Augmentation pipeline (Functions 1 & 2).
Contents
Clean images: data/train/chunk_*
Augmented images: data/aug/chunk_*
Metadata: metadata/es_clean.csv (+ aug if available)
Each CSV has columns path, text, lang, split.
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.agentic-span-states_thoughtagentic-pcap
Agentic LLMs for Network Attack Detection
Can a tool-using LLM agent independently investigate a raw network capture
and identify an attack?
How do model choice, reasoning configuration, prompt guidance, and
background traffic change the investigation and its final decision?
This Hugging Face repository is the canonical research release for our
experimental evaluation of autonomous network-traffic investigation with
agentic LLMs. It brings together the dataset, recorded… See the full description on the dataset page: https://huggingface.co/datasets/maureille/agentic-pcap.PortBench-RawData
PortBench-RawData
This repository contains the raw collected data and preprocessed asset files for PortBench.
The data spans 2015–2025 across six heterogeneous asset classes: Equities, Bonds, Commodities, Real Estate, Cryptocurrency, and Cash.
Repository Structure
PortBench-RawData/
├── raw_data/ # Raw collected data (~4.6 GB)
│ ├── fred/ # FRED macroeconomic indicators (60 series)
│ │ ├── bonds/… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-RawData.agentic_vbench_video_repairAgentic-Multi-SWE-RLAgentic-MME
Agentic-MME Dataset
This is the official dataset for the Agentic-MME benchmark, featured in Hugging Face Daily Papers.
Agentic-MME is a comprehensive benchmark designed to evaluate the abilities of multimodal agents in tool-use, web searching, and multi-step reasoning through visual clues.
Usage
You can load the dataset using the Hugging Face datasets library:
from datasets import load_dataset
dataset = load_dataset("Crystal1047/Agentic-MME")
# To see the first record… See the full description on the dataset page: https://huggingface.co/datasets/Agentic-MME/Agentic-MME.EMU-Agentic-PostTrain-Datasol-max-data
sol-max-data
Training data built by the AgentPTB arm for cell sol-max — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max.h*, and the companion to the run record in agentic-ptb/sol-max-record.
field
value
plot cell
sol-max
driver
Codex / gpt-5.6-sol
reasoning effort
max
total size
114.47 GB… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-data.Agentic-SLS-ASTM
Agentic-SLS-ASTM
ASTM mechanical-test specimens (D638 tensile, D790 flex) printed on the Inova Mk1 SLS printer and pulled on an MTS / TestWorks Instron. Each row is a single specimen with full geometry, scalar results, stress–strain + raw DAQ curves, and — for SLS rows — FK references and an embedded snapshot of the upstream print profile from ppak10/Agentic-SLS-Database.
Rows are self-contained for ML use: the full PrintProfile JSON is inlined, so features (material/energy… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-ASTM.agentic_vbench_video_repurpose
