datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seecop-osint-tool-selection
SEECOP OSINT Tool Selection
Target and intent → the right allowlisted OSINT tool, its arguments and its safety class (passive vs gated), over the SEECOP 100-tool catalogue.
What this is
Part of the SEECOP dataset family (osint-tool-selection). Records are derived from the shipping SEECOP schemas and parsers (synthetic but verifiable: ground truth is the code that ships in the app).
Ground truth
The label for every record is produced by the following… See the full description on the dataset page: https://huggingface.co/datasets/MRCB80/seecop-osint-tool-selection.tool-selection-quality-benchmark
Tool Selection Quality Benchmark
A benchmark for evaluating whether an LLM correctly judges the quality of a
tool call / function call made by another model - i.e. given a user
request, the tools available, and the model's resulting function call (or
direct reply), did the model pick the right tool and fill it in correctly?
Each row is one turn to be judged: a message history ending in either a
function call or a direct assistant response, paired with the set of tools
that were… See the full description on the dataset page: https://huggingface.co/datasets/qualifire/tool-selection-quality-benchmark.Pet-Service-Booking-Tool-Selection-Trajectory-Dataset
Pet Service Booking Tool Selection Trajectory Dataset
This dataset captures interaction trajectories from users requesting pet services through booking completion, covering service search, retrieval of pet and service details, availability checks, and booking submission. Records show how tool choices and call sequences vary by request, alongside the original trace, structured booking details, tool-call sequence, and booking result. It is designed for building and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pet-Service-Booking-Tool-Selection-Trajectory-Dataset.ai-tool-selection-goal-coherence-risk-v0.1What this repo is for
Detect when an AI system selects the wrong tool.
Core failure modes:
uses tools when not needed
avoids tools when needed
picks a tool that cannot solve the task
picks a tool that increases risk
This matters most for agentic systems.
novel-tool-selection
novel-tool-selection
Dataset generated with DeepFabric.
tool-selectionTool_Selection_Disambiguation
🇰🇿 Kazakh Tool Selection and Disambiguation Dataset
Dataset Summary
Kazakh Tool Selection and Disambiguation Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI scenarios that require choosing the most appropriate tool from multiple available options.
The dataset focuses on tool-selection reasoning, where the assistant must understand the user’s intent, compare available tools, avoid unnecessary… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Tool_Selection_Disambiguation.tool-selection-architecture-resultsphi2-tool-selectiontool-selection-accuracy-evalckan-tool-selection
