datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tooltalk-samples
ToolTalk Samples - High Quality Duplex Speech and Tool-calling in Customer Service Domain
Two people improvise realistic customer-service calls while one operates a live, stateful tool environment—with synchronized speaker-separated audio, tool calls, and outcomes.
▶ Listen to Clean · ▶ Listen to Noisy · Discuss the full dataset
In this sample: 26 calls · 80.8 minutes · 7 sample domains · 201 tool calls
Technical specs: 48 kHz / 32-bit PCM speaker-separated source… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/tooltalk-samples.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.1.0
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
72 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.jam-actions-v1-probe
jam-actions-v1-probe
Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only ·
Companion to: jam-actions-v1
Why it exists
An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes.
Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that
corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.beats-with-different-tools-and-materialslfm2-tool-aware-dataset-v1
LFM2-Tool-Aware Dataset (v1)
Synthetic speech dataset for fine-tuning LFM2.5-Audio-class audio LLMs to be tool-aware: read a tool list from the system prompt, acknowledge briefly when a user query matches a listed tool, refuse politely when no tool covers the query, and otherwise behave as a normal conversational model.
Used to train matbee/lfm2.5-audio-tool-aware-v1 (96.6% accuracy on the held-out eval split).
What this teaches a model
Voice-assistant systems often… See the full description on the dataset page: https://huggingface.co/datasets/matbee/lfm2-tool-aware-dataset-v1.lfm2-tool-aware-dataset-v2
LFM2-Tool-Aware Dataset (v2)
Synthetic speech dataset for fine-tuning LFM2.5-Audio-class audio LLMs to handle both turns of a tool-augmented voice flow: acknowledge briefly on turn 1, then narrate the dispatcher's result on turn 2 after the coordinator injects it via set_context().
Used to train matbee/lfm2.5-audio-tool-aware-v2 (~97% accuracy on the eval split, including the new turn-2 narration class).
What's new in v2
The v1 dataset taught the model to ack-and-stop… See the full description on the dataset page: https://huggingface.co/datasets/matbee/lfm2-tool-aware-dataset-v2.tau2-voice-grok-tool-mentor-trajectoriesmoshi-tool-audio
Moshi Tool-Calling — Audio-Grounded Dataset
Audio-grounded data teaching Moshi / PersonaPlex to emit tool-call special
tokens in its inner monologue when it hears a request — and to stay quiet
otherwise (listening/idle frames are trained to PAD).
Each row is a code tensor codes[17, T] at 12.5 Hz:
rows
stream
content
0
text monologue
PAD while listening/idle, `<
1:9
Moshi audio
silence
9:17
user audio
the spoken question (edge-tts), Mimi-encoded
mask=1 marks… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/moshi-tool-audio.speech-massive-vie-tool-callingsaskia001_alldata_datacreation_tool_upto_DEC_04saskia001_alldata_datacreation_tool_upto_DEC_20
