datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/lockon/ToolACE.ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/Team-ACE/ToolACE.glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en.
python-toolcallsLogs from run_python_code tool used for benchmarking.
toolcalldiscover-toolssynthetic-math-toolcall-deception
Synthetic Math Tool-Call Deception
200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception
detectors on mid-trajectory tool-call misreporting.
Each trajectory: a system prompt instructs the model to compute via an execute_python
tool under a stated tool-call limit, and requires every call to carry a running
call_index argument (1 for the first call, 2 for the second, …). The platform enforcing
the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.open-domain-ranks
Linkheft Open Domain Ranks: free domain authority data for 10.3 million domains
Score your own domain list on the latest release, by API: $5 for 25,000 lookups. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (the latest Common Crawl + Majestic release for 10.3M domains), last updated 2026-10-03.
Linkheft… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/open-domain-ranks.ToolBenchtool_callingvideo-studio-toolsToolBench
Dataset Card for "ToolBench"
More Information needed
hermes_reasoning_tool_use
TL;DR
51 004 ShareGPT conversations that teach LLMs when, how and whether to call tools.Built with the Nous Research Atropos RL stack in Atropos using a custom MultiTurnToolCallingEnv, and aligned with BFCL v3 evaluation scenarios.Released by @interstellarninja under Apache-2.0.
1 Dataset Highlights
Count
Split
Scenarios covered
Size
51 004
train
single-turn · multi-turn · multi-step · relevance
392 MB
Each row: OpenAI-style conversations… See the full description on the dataset page: https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use.Toolpacks-Snapshots
Toolpacks-Snapshots
This repo is to take periodic snapshots of all the artefacts 📦 in Toolpacks: bin.ajam.dev
Toolpacks is the Largest Collection of Multi-Platform (Android|Linux|Windows) Pre-Compiled (+ UPXed) Static Binaries (incl. Build Scripts
The Sync Workflow actions are at: https://github.com/Azathothas/Toolpacks-Snapshots-Actions
PKG Managers
!#Simply point to this:
[+] ROOT… See the full description on the dataset page: https://huggingface.co/datasets/Azathothas/Toolpacks-Snapshots.Annoy-PyEdu-Rs-Raw
Annoy: This should be a paper Title
📑 Paper | 🌐 Project Page | 💾 Released Resources | 📦 Repo
We release the raw data for our processed PythonEdu-Rs dataset, adopted from the original dataset from HuggingFaceTB team.
The data format for each line in the 0_368500_filtered_v2_ds25.sced.jsonl is as follows:
{
"problem_description": <the problem description of the function>… See the full description on the dataset page: https://huggingface.co/datasets/toolathlon-2556/Annoy-PyEdu-Rs-Raw.Nexus-Agents-ToolCalling
Nexus Agents — Tool-Calling Conversations
Synthetic, schema-verified tool-calling conversations for training the Nexus Projects
agents. This is the exact data behind
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF),
including the verification transcripts that scored it (27/27 on the behavioral
interview eval, vs 13/27 for the base model).
Links: the fine-tuned model →
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) ·
the generator + seed data + eval harness →
Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.ToolRet-Queries🔧 Retrieving useful tools from a large-scale toolset is an important step for Large language model (LLMs) in tool learning. This project (ToolRet) contribute to (i) the first comprehensive tool retrieval benchmark to systematically evaluate existing information retrieval (IR) models on tool retrieval tasks; and (ii) a large-scale training dataset to optimize the expertise of IR models on this tool retrieval task.
See the official Github for more details.
A concrete example for our evaluation… See the full description on the dataset page: https://huggingface.co/datasets/mangopy/ToolRet-Queries.ToolMind
ToolMind: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
ToolMind is a large-scale, high-quality tool-agentic dataset with 160k synthetic data instances generated using over 20k tools and 200k augmented open-source data instances.
Our data synthesis pipeline first constructs a function graph based on parameter correlations and then uses a multi-agent framework to simulate realistic user–assistant–tool interactions.
Beyond trajectory-level validation, we employ fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/Nanbeige/ToolMind.Annoy-PyEdu-Rs
Annoy: This should be a paper Title
📑 Paper | 🌐 Project Page | 💾 Released Resources | 📦 Repo
This is the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block.
Dataset
Dataset
Link
Annoy-PythonEdu-Rs
🤗
Please also check the raw data after our processing… See the full description on the dataset page: https://huggingface.co/datasets/toolathlon-2556/Annoy-PyEdu-Rs.G1_Dex1_ToolboxStorageToolScale
ToolScale Dataset
The ToolScale dataset is a key component of the ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestrationproject. It provides synthetic environment and tool-call tasks specifically generated to aid the reinforcement learning (RL) training of small orchestrator models. These orchestrators are designed to effectively manage and coordinate diverse intelligent tools and other models for solving complex, multi-turn agentic tasks.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/ToolScale.Qwen-Terminal-ToolBench-Processed-Tokenized
Qwen Terminal ToolBench Processed Datasets
Qwen-family processed/template-applied and selected tokenized terminal datasets.
Contents
qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text
qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text
qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels
qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.Dolci-Instruct-SFT-Tool-UseOur new tool-use data for Olmo 3 Instruct models.
For the full dataset, documentation, etc. see the main dataset card.
This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Citation
@misc{olmo2025olmo3,
title={Olmo 3},
author={Team Olmo and Allyson Ettinger and Amanda Bertsch and Bailey Kuehl and David Graham and David Heineman and Dirk Groeneveld and Faeze Brahman and Finbarr Timbers and Hamish… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT-Tool-Use.toolace-parsed
[PARSED] ToolACE
The data in this dataset is a subset of the original Team-ACE/ToolACE
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
toolace
yes
yes
yes
complex
11k
This is a re-parsing formatting dataset for the ToolACE official dataset.
Load the dataset
from datasets import load_dataset
ds = load_dataset("minpeter/toolace-parsed")
print(ds)
# DatasetDict({
# train: Dataset({
# features:… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/toolace-parsed.atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.ToolRet-Tools🔧 Retrieving useful tools from a large-scale toolset is an important step for Large language model (LLMs) in tool learning. This project (ToolRet) contribute to (i) the first comprehensive tool retrieval benchmark to systematically evaluate existing information retrieval (IR) models on tool retrieval tasks; and (ii) a large-scale training dataset to optimize the expertise of IR models on this tool retrieval task.
This ToolRet-Tools contains the toolset corpus of our tool retrieval benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/mangopy/ToolRet-Tools.toolace_hermes_tool_useVC-Tooler-SFT
VC-Tooler-SFT
Supervised cold-start trajectories for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use.
🔗 Links
📄 Paper: arXiv
🌐 Project Page: w1zheng.github.io/VC-Tooler
🤗 Hugging Face: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
🧩 ModelScope: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
This dataset is the Stage I (supervised fine-tuning) trajectory bank used to teach a
vision–language model to use visual tools as a compositional and adaptive… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-SFT.tool-calls-singleturnToolHop
[ACL 2025] ToolHop
[ACL 2025] ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
Data for the paper ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
Junjie Ye
jjye23@m.fudan.edu.cn
Jan. 07, 2025
Introduction
Effective evaluation of multi-hop tool use is critical for analyzing the understanding, reasoning, and function-calling capabilities of large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/bytedance-research/ToolHop.
