datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hermes_reasoning_tool_use
TL;DR
51 004 ShareGPT conversations that teach LLMs when, how and whether to call tools.Built with the Nous Research Atropos RL stack in Atropos using a custom MultiTurnToolCallingEnv, and aligned with BFCL v3 evaluation scenarios.Released by @interstellarninja under Apache-2.0.
1 Dataset Highlights
Count
Split
Scenarios covered
Size
51 004
train
single-turn · multi-turn · multi-step · relevance
392 MB
Each row: OpenAI-style conversations… See the full description on the dataset page: https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use.interstellarxlam_hermes_validatedtoolace_hermes_tool_usetool-calls-singleturnfrontier-synth-dialogue-llmhermes-function-calling-v1
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/interstellarninja/hermes-function-calling-v1.interstellartool-use-multiturn-reasoninghermes_salesforce_apigen_tool_usetool-calls-multiturnhermes_interleaved_reasoning_tool_usesalesforce_hermes_toolstool-calls-single-reasoningtool-calls-sharegptglaive-function-calling-5kjson-mode-singleturnGaokao-LLM-data
QZDH_Gaokao_Data: Gaokao Past Paper Reasoning Dataset
Chinese readme link is here: 简体中文
QZDH_Gaokao_Data is a dataset independently collected by the Qizhi Navigation Project, aimed at promoting the rapid development of AI education, assisting in the development of AI applications, and the construction of AI teacher models. The original intention of the team in building this dataset is to provide data support for the fine-tuning of models used by the team, and it is hoped that… See the full description on the dataset page: https://huggingface.co/datasets/Interstellar174/Gaokao-LLM-data.json-mode-reasoningnvidia_hermes_when2calljson-mode-agenticjson-mode-verifiabletool-use-relevance-reasoningtool-calls-dpotool-calls-rawtoolace_sequential_tool_use_reasoningreasoning-sft-interstellarninja-json-mode-reasoning-160K
json-mode-reasoning (converted)
Converted version of interstellarninja/json-mode-reasoning, filtered to 20,474 rows with valid <think> reasoning traces.
Format
Each row has three columns:
input — list of dicts [{"role": "system/user", "content": "..."}, ...] (conversation turns ending on the last user turn, includes system prompt with JSON schema)
response — assistant response string with <think> reasoning block followed by JSON output
source — fixed as… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-interstellarninja-json-mode-reasoning-160K.interleaved_tool_use_reasoningjson-schema-store-reasoningjson-schema-store-rl
