Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openbmb /UltraData-SFT-Agent-2609 UltraData-SFT-Agent-2609 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series English | 中文 📚 Introduction UltraData-SFT-Agent-2609 is the L3 refined data for Agent instruction-tuning within UltraData's L0-L4 tiered data management framework. Built for the post-training of MiniCPM5-2B, it complements UltraData-SFT-2605 (core-domain SFT) with executable Agent trajectories. The release contains approximately 500,000 samples spanning tool use… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609.texttext-generation100K<n<1M292 likes27k downloads1mo agoHugging Face02neulab /agent-data-collection Agent Data Collection A comprehensive collection of agent interaction datasets for training and evaluating AI agents across diverse domains and tasks. This dataset aggregates high-quality agent trajectories from various environments including web browsing, code generation, household tasks, knowledge base querying, and software engineering. The dataset is collected through methods described in Agent Data Protocol. Dataset Splits Each dataset configuration provides up… See the full description on the dataset page: https://huggingface.co/datasets/neulab/agent-data-collection.text-generation1M<n<10M116 likes11k downloads7mo agoHugging Face03skeole /qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols. ~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks. The only human artifacts are: agents/* human/* AGENTS.md texttext-generation1K<n<10K3 likes8.7k downloads20d agoHugging Face04nvidia /Nemotron-SFT-Agentic-v2 Dataset Description The Nemotron-SFT-Agentic-v2 dataset is a collection of synthetic single-turn and multi-turn tool-use trajectories designed to strengthen models’ capabilities as interactive, tool-using agents. It targets tasks where the model must decompose user goals, decide when to call tools, and reason over tool outputs to complete tasks reliably and safely. This dataset is ready for commercial use. The dataset consolidates three internally curated components (described… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Agentic-v2.text-generation86 likes8.4k downloads2mo agoHugging Face05agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.1k downloads2y agoHugging Face06open-thoughts /AgentTrove AgentTrove AgentTrove is the largest open-source collection of agentic interaction traces to date, released by the OpenThoughts-Agent team. It contains 1,696,847 rows drawn from 219 source datasets spanning code repair, shell scripting, mathematical problem-solving, competitive programming, and general computer-use tasks. At 1.7 million rows, AgentTrove is 4× the size of the Nemotron Terminal Corpus (430 K rows), the previous largest open-source agentic trace dataset.… See the full description on the dataset page: https://huggingface.co/datasets/open-thoughts/AgentTrove.texttext-generation1M<n<10M200 likes4.3k downloads5mo agoHugging Face07TeichAI /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K108 likes4.3k downloads5mo agoHugging Face08yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M53 likes4.2k downloads7mo agoHugging Face09agentlans /text-sft-questions-answers-only text-sft: Questions and Answers This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft. Overview The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.texttext-generation100K<n<1M2 likes4k downloads11mo agoHugging Face10SciPhi /AgentSearch-V1 Getting Started The AgentSearch-V1 dataset boasts a comprehensive collection of over one billion embeddings, produced using jina-v2-base. The dataset encompasses more than 50 million high-quality documents and over 1 billion passages, covering a vast range of content from sources such as Arxiv, Wikipedia, Project Gutenberg, and includes carefully filtered Creative Commons (CC) data. Our team is dedicated to continuously expanding and enhancing this corpus to improve the search… See the full description on the dataset page: https://huggingface.co/datasets/SciPhi/AgentSearch-V1.texttext-generation10K<n<100K92 likes3.7k downloads3y agoHugging Face11agentlans /DSULT-Core-ShareGPT-X DSULT-Core/ShareGPT-X Filtered Dataset This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines. The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.tabulartext-generation10K<n<100K3 likes3.5k downloads10mo agoHugging Face12AgenticCommons /formal-math-autoformalization Formal Math Autoformalization Dataset A growing, CC0 public-domain corpus of ⟨natural-language statement ↔ Lean 4 statement + proof⟩ pairs, contributed through the Agentic Commons network. Why this is scarce data. Mathlib already contains millions of proven Lean theorems — but as bare Lean, with no paired natural language: theorem add_comm (a b : ℕ) : a + b = b + a := ... -- no "addition on naturals is commutative" attached The scarce, valuable artifact is the pairing of the… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/formal-math-autoformalization.texttext-generation1K<n<10K3 likes2.9k downloads11d agoHugging Face13mikuhhn1239 /novel-agent-sft-dataset All Novel Can Be Galgame — 完整数据集 中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。 项目地址:https://github.com/lin1753/novel2galgame 训练代码仓库:https://github.com/lin1753/novel-agent 数据规模 目录 文件数 大小 说明 training/ 52 689 MB 训练用 SFT 数据 (JSONL) raw-books/ 671 327 MB 669 本原始小说 processed/ 39,842 1.2 GB 按章节预处理文本 annotations/ 1,626 1 MB 原始标注文件 合计 42,191 2.2 GB 目录结构 datasets/ ├── training/ │ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.texttext-generation10K<n<100K7 likes2.9k downloads3mo agoHugging Face14Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K2 likes2.6k downloads3mo agoHugging Face15AI-Secure /DTap-Bench-Agent-Trajectories DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the full collection of the agent trajectories produced from evaluating the DTap-Bench from DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task ships the configuration the evaluator needs to spin up the… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DTap-Bench-Agent-Trajectories.text-generation1K<n<10K4 likes2.6k downloads3mo agoHugging Face16whfeLingYu /Unified_Agent_Framework A Unified Framework for the Evaluation of LLM Agentic Capabilities This repository contains the dataset (Benchmark, Toolkit, and Environment assets) for the paper A Unified Framework for the Evaluation of LLM Agentic Capabilities. The official code and agent execution sandbox can be found on GitHub: whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities. Dataset Description The dataset integrates diverse agent benchmarks into a standardized… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Unified_Agent_Framework.text-generation0 likes2.5k downloads1mo agoHugging Face17NexusProjectsAI /Nexus-Agents-ToolCalling Nexus Agents — Tool-Calling Conversations Synthetic, schema-verified tool-calling conversations for training the Nexus Projects agents. This is the exact data behind Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF), including the verification transcripts that scored it (27/27 on the behavioral interview eval, vs 13/27 for the base model). Links: the fine-tuned model → Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) · the generator + seed data + eval harness → Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.texttext-generation100K<n<1M2 likes2.4k downloads4mo agoHugging Face18AgentPublic /open_government Open Government Dataset Open Government is the largest agregation of governement text and data made available as part of open data programs. In total, the dataset contains approximately 380B tokens. While Open Government aims to become a global resource, in its current state it mostly features open datasets from the US, France, European and international organizations. The dataset comprises 16 collections curated through two different initiaties: Finance commons and Legal commons.… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/open_government.tabulartext-generation10M<n<100M5 likes2.4k downloads2y agoHugging Face19MaxDevv /real-pi-coding-agent-traces-sessions Real Pi Coding Agent Traces Sessions An aggregated dataset of real human–AI coding agent sessions, collected from 21 independently published Hugging Face datasets and hand-filtered to exclude synthetic or AI-generated content. Every session is an unedited (but redacted) trace of a real person using pi — an open-source AI coding agent harness — to build, debug, and ship real open-source software. Real prompts, real tool calls, real errors, real backtracking. Why this… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/real-pi-coding-agent-traces-sessions.text-generation1K<n<10K5 likes2.3k downloads3mo agoHugging Face20sammshen /lmcache-agentic-traces LMCache Agentic Dataset Collection A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache. Motivation Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.tabulartext-generation10K<n<100K17 likes2.1k downloads4mo agoHugging Face21yatin-superintelligence /White-Hat-Security-Agent-Prompts-600K White Hat Security Agent Prompts 600K Overview The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios. Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.texttext-generation100K<n<1M21 likes2.1k downloads7mo agoHugging Face22nvidia /Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Dataset Description: Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection (IPI) attacks hidden inside tool-returned environment data. In each record, the agent receives a benign user request that requires calling a read tool whose output contains an adversarial instruction disguised as legitimate domain content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.textreinforcement-learning1K<n<10K9 likes2.1k downloads4mo agoHugging Face23AletheiaResearch /GLM-5.2-AgentThis dataset was generated using teich by TeichAI GLM-5.2 Agent traces This directory contains raw agent trace files generated by teich. JSONL files: 319 Model metadata: glm-5.2 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GLM-5.2-Agent.tabulartext-generationn<1K62 likes2k downloads3mo agoHugging Face24jupyter-agent /jupyter-agent-dataset Jupyter Agent Dataset Dataset Details Dataset Description The dataset uses real Kaggle notebooks processed through a multi-stage pipeline to de-duplicate, fetch referenced datasets, score educational quality, filter to data-analysis–relevant content, generate dataset-grounded question–answer (QA) pairs, and produce executable reasoning traces by running notebooks. The resulting examples include natural questions about a dataset/notebook, verified answers, and… See the full description on the dataset page: https://huggingface.co/datasets/jupyter-agent/jupyter-agent-dataset.textquestion-answering10K<n<100K174 likes2k downloads1y agoHugging Face25AI-Secure /DecodingTrust-Agent-Platform DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the per-task dataset for the DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task ships the configuration the evaluator needs to spin up the sandbox, run an agent, and verify the outcome — config.yaml (task… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust-Agent-Platform.text-generation1K<n<10K1 likes1.5k downloads4mo agoHugging Face26trace-commons /agent-traces Trace Commons — Agent Traces Trace Commons is one open, public dataset of coding-agent sessions — the back-and-forth between a developer and an AI coding agent, including prompts, model responses, tool calls, and command output — contributed voluntarily as an open resource for studying, evaluating, and building on how these agents actually work. Every trace here was donated only from a public, open-source repository, was anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.tabulartext-generationn<1K36 likes1.5k downloads4mo agoHugging Face27nvidia /Nemotron-RL-Agentic-Terminal-Pivot-v1 Dataset Description The Nemotron-RL-Agentic-Terminal-Pivot-v1 dataset provides training samples for reinforcement learning of command-line ("terminal use") LLM agents with the terminus_judge environment in NeMo Gym. Each record is a single agent decision point extracted from a successful agent trajectory on a terminal task: responses_create_params.input — the prompt: the task instruction plus the terminal interaction history (prior agent actions and terminal outputs) up to the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1.texttext-generation10K<n<100K31 likes1.5k downloads1mo agoHugging Face28agentlans /li2017dailydialog DailyDialog in ShareGPT-like Format Dataset Summary This dataset is a reformatted version of the DailyDialog dataset, structured to resemble ShareGPT conversations. Each row represents a single conversation, with added features and adjustments to better support conversational AI training. The dataset includes: A random system prompt to initiate natural interactions. Emotion annotations for user or AI messages. Corrected message spacing using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/li2017dailydialog.texttext-generation10K<n<100K1 likes1.4k downloads2y agoHugging Face29agentlans /train-of-thought Train of Thought Dataset Overview This dataset readapts agentlans/think-more into the Alpaca-style instruction tuning format for training language models in direct answering and chain-of-thought reasoning. Dataset Structure Each original example was randomly assigned to be thinking on or off: Thinking off: Outputs only the final answer. Thinking on: Outputs a chain-of-thought (CoT) reasoning process wrapped in <think>...</think>, followed by the final answer… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/train-of-thought.texttext-generation1M<n<10M5 likes1.4k downloads1y agoHugging Face30antiquality /agentabstain AgentAbstain: Do LLM Agents Know When Not to Act? AgentAbstain is a paired-task benchmark for agentic abstention: the calibrated ability of tool-using LLM agents to recognize when not to act. It contains 263 task pairs across 42 executable MCP sandbox environments, built on an agent-native taxonomy of 8 abstention scenarios. Every should-act task ships with a should-abstain variant that differs by a single controlled perturbation to the instruction, the… See the full description on the dataset page: https://huggingface.co/datasets/antiquality/agentabstain.texttext-generationn<1K1 likes1.4k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.