datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
discover-toolsatlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.us-zip-codes-demographics
Ziplore US ZIP Codes: city, county, coordinates, time zone and Census demographics
Look up any ZIP code's demographics by API instead of loading the file: $5 for 12,000 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (Census ACS 2020-2024 figures for every US ZIP), last updated 2026-09-24.
Ziplore ZIP… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/us-zip-codes-demographics.stillshipping-tools
StillShipping: maintenance verdict for every tracked agent tool
One row per tracked AI agent tool, with the nightly maintenance verdict (maintained, slowing or dead), the 0-100 freshness behind it, and every GitHub signal the verdict was computed from.
Rows in this cut
362
One row is
one tool
Cut
2026-10-01
Refreshed
Monthly, on the first of the month
Measured by
StillShipping
Method
https://toolproof.thecompound.tech/methodology
Licence
Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/stillshipping-tools.orbitwren-orbital-compute-tracker
Orbitwren: Orbital Compute & Space Data Center Tracker
State of Orbital Data Centers, Oct 2026: every project, launch calendar and partnership map, all sources cited. 16-page PDF, $19 (free 2-page sample). Buy now → · Download link on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (today's orbital compute tracker, upcoming… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/orbitwren-orbital-compute-tracker.ogpo-flow-toolhang-umap-ckpts
OGPO (flow policy) tool_hang checkpoints for the UMAP analysis
Checkpoints of one online-RL run of the public OGPO code (https://github.com/simchowitzlabpublic/OGPO_public),
scripts/ogpo/toolhang.sh defaults: robomimic tool_hang-ph-low_dim, flow-matching policy (10 flow steps, horizon 8),
conservative group advantages (OGPO-CA), 10-head Q ensemble, seed 1.
Wandb: https://wandb.ai/pluralistic-goal-conditioning/OGPO/runs/kshv94if
file
phase
note
params_200000.pkl
end of… See the full description on the dataset page: https://huggingface.co/datasets/shashwatsaxena136/ogpo-flow-toolhang-umap-ckpts.oceania-gov-open-data-catalog
Oceania Government Open Data — Combined Catalogue (hourly snapshot)
Combined regional catalogue of Oceania (Australia + New Zealand) public-service open
data harvested from both data.gov.au and data.govt.nz portals, including state,
territory and local-council publishers.
License declaration
License: other (see below). Records in this catalogue inherit the licence of their
source dataset. Where the source declares a standard open licence the record is tagged
with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.data-govt-nz-mirror
data.govt.nz — Mirror Catalogue (hourly snapshot)
Mirror publication of the New Zealand open government data catalogue (data.govt.nz).
Each row of catalog.csv is a dataset record as harvested from the data.govt.nz
CKAN instance (national agencies and local councils).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-govt-nz-mirror.tool-response-injections
Tool Response Injection Dataset
A classification dataset for evaluating tool response prompt injection detection, specifically attempting to emulate realistic injection vectors within tool-calling workflows.
Construction
Tool responses are taken from interstellarninja/tool-calls-single-reasoning and combined with
a random prompt injection from neuralchemy/Prompt-injection-dataset. The prompt injection is inserted into a random
field of the tool response. The… See the full description on the dataset page: https://huggingface.co/datasets/rgeada/tool-response-injections.data-tool-momentum
Datamata Data Tool Momentum Index
Cross-signal momentum for open source data tools: GitHub stars, forks and 4-week star growth, PyPI and npm downloads, and active job demand. One row per tool from the most recent weekly snapshot, with a 0-100 momentum score.
Latest snapshot: 2026-10-04
Tools in this release: 26
Updated: weekly
Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution.
Source & methodology:… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/data-tool-momentum.tooldrift-app-rankings
ToolDrift: OpenRouter app usage rankings, captured daily
One row per app per ranking window per capture, with the tool it maps to where ToolDrift tracks one. It is the same series as the model rankings, read from the consumer side.
Rows in this cut
1,181
One row is
one app in one ranking window on one capture day
Cut
2026-10-01
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method
https://toolproof.thecompound.tech/methodology
Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-app-rankings.tooldrift-model-rankings
ToolDrift: OpenRouter model usage rankings, captured daily
One row per model per ranking window per capture: its rank, the tokens and requests behind that rank, and its share of the window. The series shows which models the market actually routes work to, day by day.
Rows in this cut
84,461
One row is
one model in one ranking window on one capture day
Cut
2026-10-01
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-model-rankings.tooldrift-tools
ToolDrift: the AI coding tools under watch
One row per AI coding tool watched nightly: its layer in the stack, its vendor, its licence and pricing model, its default model, and the GitHub maintenance signals beside the OpenRouter usage rank.
Rows in this cut
36
One row is
one tool
Cut
2026-10-01
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method
https://toolproof.thecompound.tech/methodology
Licence
Creative Commons Attribution 4.0… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-tools.ai-tools-radar
Mirror. The canonical, citable record is on Zenodo: series DOI 10.5281/zenodo.22730450 (always the latest edition). Files here are the edition 2026-09, DOI 10.5281/zenodo.23075940. The earlier edition 2026-09-13 is DOI 10.5281/zenodo.22730451 and stays in this repository history. Please cite the Zenodo DOI. Project pages: https://arsentev.ai/github
AI Tools Radar — GitHub and Hugging Face projects selected by arsentev.ai
Edition 2026-09 · built 2026-10-01 · author Evgenii… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/ai-tools-radar.MapSatisfyBench-MockData-ToolsThe datasets associated with MapSatisfyBench include the benchmark data file "MapSatisfyBench_Benchmark.csv" and the mock data (other .csv files) used by the sandbox tools during simulation execution.
toolproof-indexes
Toolproof: the nine indexes and what each one currently measures
One row per index under the Toolproof masthead: what it measures, the method behind it, the public endpoint its figures come from, and the headline figure that endpoint returned at the moment of the cut.
Rows in this cut
9
One row is
one index
Cut
2026-10-01
Refreshed
Monthly, on the first of the month
Measured by
Toolproof
Method
https://toolproof.thecompound.tech/methodology
Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/toolproof-indexes.ai-tool-prompts
AI Tool Prompts
A small synthetic dataset of 200 English user instructions designed for experiments with
intent classification, routing, AI tool selection, and lightweight text classification.
The dataset contains 10 balanced categories with 20 examples each.
Dataset Structure
Each row contains:
Column
Description
id
Unique example identifier
category
Target intent/category
instruction
Synthetic user instruction
expected_output_type
General type… See the full description on the dataset page: https://huggingface.co/datasets/ostwestfale/ai-tool-prompts.asia-health-cost-2024-consolidated
Asia Health Cost 2024 — Consolidated
Consolidated 2024 fiscal-year medical operations & cost dataset for a pan-Asia healthcare enterprise
(China / Japan / India). Created by merging three country-level, de-identified source datasets and
normalising every cost to USD.
Namespace note: the task referenced the source/output under the medi-core namespace, which is not
accessible with the current credentials. The identical pipeline was executed under the toolathon123
namespace:… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/asia-health-cost-2024-consolidated.to-tool-call-papers
📚 To-Tool-Call Papers
A curated paper library for LLM tool use, function calling, agent training, and environment synthesis
To-Tool-Call Papers is a bilingual research library for tracking papers on tool use, function calling, agent data synthesis, environment scaling, agentic RL, and tool-use benchmarks.
Quick Start ·
At a Glance ·
Files ·
Schema ·
Copyright
[!IMPORTANT]
This dataset is a research reading collection… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-papers.toolverifier
TOOLVERIFIER: Generalization to New Tools via Self-Verification
This repository contains the ToolSelect dataset which was used to fine-tune Llama-2 70B for tool selection.
Data
ToolSelect data is synthetic training data generated for tool selection task using Llama-2 70B and Llama-2-Chat-70B.
It consists of 555 samples corresponding to 173 tools.
Each training sample is composed of a user instruction, a candidate set of tools that includes the
ground truth tool, and a… See the full description on the dataset page: https://huggingface.co/datasets/facebook/toolverifier.ai-tool-prompts-mini
AI Tool Prompts Mini
A tiny synthetic dataset for experimenting with tool routing and text classification.
Columns
id — row identifier
prompt — user request
tool — expected tool category
Labels
search
calculator
weather
translation
summarize
code
email
calendar
Intended use
This dataset is designed for:
Hugging Face demos
text-classification experiments
tool-routing prototypes
educational projects
All examples are synthetic and… See the full description on the dataset page: https://huggingface.co/datasets/Beratung/ai-tool-prompts-mini.tool-callsTool calling master dataset
Contains the following:
Query -> Available tools (name + description + schema) -> Tool name
Sources (identified by source column):
subsets of existing tool-calling dataset sources parsed into the above format
synthetic data
Will be parsed into the following two passes:
Query -> List of tool names + descriptions -> Tool name
Tool name + tool schema -> Tool call
freqtrade_ml_toolkit
freqtrade ml_toolkit — shared datasets
Datasets for the ml_toolkit framework in the freqtrade repo. Nothing here is
hand-made: everything is produced by the scripts below and can be regenerated.
Layout
ohlcv/bybit/<PAIR>_1h.feather raw OHLCV, laid out like a --datadir
kronos/bybit_1h/csv/<PAIR>_1h.csv exported fine-tuning corpus
kronos/bybit_1h/csv/split.json train/val boundaries for that corpus
kronos/bybit_1h/*.tar.gz the same corpus… See the full description on the dataset page: https://huggingface.co/datasets/l2533584225/freqtrade_ml_toolkit.lab06-tool-calling
Order the Lab: a tool-calling trace, and what the model said with no tools
Completed by James Seegel for MIS 752 at UNLV.
Based on Dr. Richard Young's teaching notebook.
My takeaways
My main trace reported path_used: native, with zero native tool-calling failures across the four sweep questions. The forced prompt-based run also succeeded with path_used: prompt-json. It used 881 prompt tokens versus 1,019 for native, or 138 fewer, with two model calls in each run.… See the full description on the dataset page: https://huggingface.co/datasets/JamesSeegel/lab06-tool-calling.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.AI-Coding-Tools
Dataset Card for 2026 AI Coding Tools
Last Updated: 24 May 2026
Curated By: Joy Larkin
Language(s) (NLP): English
License: MIT
Repository: https://github.com/joylarkin/AI-Coding-Landscape
Blog: https://cleverhack.com/ai-coding-landscape
Dataset Description
CSV file of AI Coding Tools released in 2026 & 2025.
mini-vlm-toolkit-data
mini-vlm-toolkit counting & grounding benchmarks
Processed evaluation splits used by
mini-vlm-toolkit.
Each <split>.tsv holds the exact prompts (question) and annotations
(answer plus metadata) we evaluate on; image_path is relative to
images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13
configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running
python run_benchmark.py --data… See the full description on the dataset page: https://huggingface.co/datasets/bryanzhou008/mini-vlm-toolkit-data.ai-tools-database-25k-aitoolbuzzdata-gov-au-mirror
data.gov.au — Mirror Catalogue (hourly snapshot)
Mirror publication of the Australian open government data catalogue (data.gov.au).
Each row of catalog.csv is a dataset record as harvested from the data.gov.au CKAN
instance (federal, state and territory agencies).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in the… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-gov-au-mirror.dino-data-vision-tooling-preview
Dino Data Vision Tooling Preview
What This Dataset Is
This dataset is a focused vision-tooling preview built from two Dino Data capability slices:
image context understanding
image tooling
The goal is to train or inspect assistant behavior for image-related tasks where visual context, multimodal interpretation, or tool-aware image handling is relevant.
Included Capability Slices
Source lane
Public task name
What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-vision-tooling-preview.
