datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/Team-ACE/ToolACE.ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/lockon/ToolACE.glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en.
python-toolcallsLogs from run_python_code tool used for benchmarking.
discover-toolstoolcallsynthetic-math-toolcall-deception
Synthetic Math Tool-Call Deception
200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception
detectors on mid-trajectory tool-call misreporting.
Each trajectory: a system prompt instructs the model to compute via an execute_python
tool under a stated tool-call limit, and requires every call to carry a running
call_index argument (1 for the first call, 2 for the second, …). The platform enforcing
the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.open-domain-ranks
Linkheft Open Domain Ranks: free domain authority data for 10.3 million domains
Score your own domain list on the latest release, by API: $5 for 25,000 lookups. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (the latest Common Crawl + Majestic release for 10.3M domains), last updated 2026-10-10.
Linkheft… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/open-domain-ranks.ToolBench
Dataset Card for "ToolBench"
More Information needed
tool_callinghermes_reasoning_tool_use
TL;DR
51 004 ShareGPT conversations that teach LLMs when, how and whether to call tools.Built with the Nous Research Atropos RL stack in Atropos using a custom MultiTurnToolCallingEnv, and aligned with BFCL v3 evaluation scenarios.Released by @interstellarninja under Apache-2.0.
1 Dataset Highlights
Count
Split
Scenarios covered
Size
51 004
train
single-turn · multi-turn · multi-step · relevance
392 MB
Each row: OpenAI-style conversations… See the full description on the dataset page: https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use.ToolRet-Queries🔧 Retrieving useful tools from a large-scale toolset is an important step for Large language model (LLMs) in tool learning. This project (ToolRet) contribute to (i) the first comprehensive tool retrieval benchmark to systematically evaluate existing information retrieval (IR) models on tool retrieval tasks; and (ii) a large-scale training dataset to optimize the expertise of IR models on this tool retrieval task.
See the official Github for more details.
A concrete example for our evaluation… See the full description on the dataset page: https://huggingface.co/datasets/mangopy/ToolRet-Queries.Nexus-Agents-ToolCalling
Nexus Agents — Tool-Calling Conversations
Synthetic, schema-verified tool-calling conversations for training the Nexus Projects
agents. This is the exact data behind
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF),
including the verification transcripts that scored it (27/27 on the behavioral
interview eval, vs 13/27 for the base model).
Links: the fine-tuned model →
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) ·
the generator + seed data + eval harness →
Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.toolace-parsed
[PARSED] ToolACE
The data in this dataset is a subset of the original Team-ACE/ToolACE
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
toolace
yes
yes
yes
complex
11k
This is a re-parsing formatting dataset for the ToolACE official dataset.
Load the dataset
from datasets import load_dataset
ds = load_dataset("minpeter/toolace-parsed")
print(ds)
# DatasetDict({
# train: Dataset({
# features:… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/toolace-parsed.ToolScale
ToolScale Dataset
The ToolScale dataset is a key component of the ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestrationproject. It provides synthetic environment and tool-call tasks specifically generated to aid the reinforcement learning (RL) training of small orchestrator models. These orchestrators are designed to effectively manage and coordinate diverse intelligent tools and other models for solving complex, multi-turn agentic tasks.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/ToolScale.Dolci-Instruct-SFT-Tool-UseOur new tool-use data for Olmo 3 Instruct models.
For the full dataset, documentation, etc. see the main dataset card.
This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Citation
@misc{olmo2025olmo3,
title={Olmo 3},
author={Team Olmo and Allyson Ettinger and Amanda Bertsch and Bailey Kuehl and David Graham and David Heineman and Dirk Groeneveld and Faeze Brahman and Finbarr Timbers and Hamish… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT-Tool-Use.ToolRet-Tools🔧 Retrieving useful tools from a large-scale toolset is an important step for Large language model (LLMs) in tool learning. This project (ToolRet) contribute to (i) the first comprehensive tool retrieval benchmark to systematically evaluate existing information retrieval (IR) models on tool retrieval tasks; and (ii) a large-scale training dataset to optimize the expertise of IR models on this tool retrieval task.
This ToolRet-Tools contains the toolset corpus of our tool retrieval benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/mangopy/ToolRet-Tools.toolace_hermes_tool_usetool-calls-singleturnVC-Tooler-SFT
VC-Tooler-SFT
Supervised cold-start trajectories for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use.
🔗 Links
📄 Paper: arXiv
🌐 Project Page: w1zheng.github.io/VC-Tooler
🤗 Hugging Face: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
🧩 ModelScope: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
This dataset is the Stage I (supervised fine-tuning) trajectory bank used to teach a
vision–language model to use visual tools as a compositional and adaptive… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-SFT.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.reason-tool-use-demo-1500
Dataset info
The dataset is a selection of reasoning toolcalls data from https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use, which contains data from Hermes-Tools、Glaive-FC、ToolAce、Nvidia-When2Call.
The format has been transformed to adapt llama-factory v1 training pipeline.
ToolRetrieval
ToolRetrievalInstruction
An MTEB dataset
Massive Text Embedding Benchmark
ToolRet in the w/ inst. setting, where each query is paired with a task instruction describing what kind of tool is required. Corresponds to Table 5 of the paper; see ToolRetrieval for the w/o inst. setting.
Task category
Retrieval (text-to-text)
Domains
Programming, Web, Written
Reference
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
Source… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ToolRetrieval.ToolACE-Qwen-cleaned
ToolACE for Qwen
Created by: Seungwoo Ryu
Introduction
This dataset is an adaptation of the ToolACE dataset, modified to be directly compatible with Qwen models for tool-calling fine-tuning.
The original dataset was not in a format that could be immediately used for tool-calling training, so we have transformed it accordingly.
This makes it more accessible for training Qwen-based models with function-calling capabilities.
This dataset is applicable to all… See the full description on the dataset page: https://huggingface.co/datasets/tryumanshow/ToolACE-Qwen-cleaned.atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.tool_calling_shuffleOpenGrad-ToolPolicy-Canonical-v2-M0-snapshot
This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture.
What this release is
OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of public tool-use and function-calling datasets. It was built to test one hypothesis with a measurement attached: that the tool-call collapse observed in the M0 SFT experiments on v1 was caused by the… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot.recallroll-vehicle-recalls
Recallroll: US vehicle safety recalls, last 12 months
Check your own VINs for open recalls, live, by API: $5 for 400 calls of up to 10 VINs. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com Updated weekly (Mondays), with every column and all history. Need today's data? Recallroll VIN + recall API has it live.
Need it fresh, filtered or via API? This free file is a snapshot… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/recallroll-vehicle-recalls.glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en.
