datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-symptom-triage-conversationalNemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.smolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.conversational-dynamics-egocom
Conversational Dynamics — EgoCom
Derived temporal annotations and model-ready training anchors for
conversational-dynamics and turn-taking research, generated from EgoCom.
This dataset is produced by the conversational-dynamics-data pipeline. The
underlying objective is to expose conversational data in a representation
suitable for temporal and action-conditioned models:
state_t + action_t → future conversational state
Contents
Three configurations are provided.… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom.conversational-s2s-v1
Conversational S2S v1 — dialogues parlés FR + EN
Corpus synthétique de conversations orales multi-tours entre un utilisateur
et un assistant, en français et en anglais, destiné au finetuning
speech-to-speech de LFM2.5-Audio (Liquid AI). Chaque tour est un clip audio
séparé, aligné avec son texte ; les dialogues sont conçus pour être rejoués tour
à tour (user → assistant → user → …).
Projet tts-model-exploration. Produit par
voxtral_datagen_pipeline (s2s-skeletons → remplissage… See the full description on the dataset page: https://huggingface.co/datasets/Rcarvalo/conversational-s2s-v1.conversational-dynamics-egocom-12.5hz
Conversational Dynamics — EgoCom (12.5 Hz)
Derived temporal annotations and model-ready training anchors for
conversational-dynamics and turn-taking research, generated from EgoCom.
This dataset is produced by the conversational-dynamics-data pipeline. The
underlying objective is to expose conversational data in a representation
suitable for temporal and action-conditioned models:
state_t + action_t → future conversational state
12.5 Hz release
The same pipeline… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom-12.5hz.gemma3n-conversational-reasoning
Gemma3N Conversational Reasoning
This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use:
from datasets import load_dataset
from unsloth.chat_templates import standardize_data_formats
dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]")
dataset = standardize_data_formats(dataset)
Schema:
conversations: ShareGPT-style list of turns with from and value
metadata columns are included for analysis and filtering
Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.TACTBench-Samples
TACTBench Demonstration Samples
This repository contains five full-context demonstration examples from
TACTBench. It does not contain the TACT training set or the remaining hidden
TACTBench evaluation set. The samples use the same full-history representation
as the benchmark evaluation and illustrate direct correction, error
explanation, guided revision, clarification checking, affective feedback, and
retry elicitation.
Data
data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.conversational-sarcasm-benchmark
Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only
A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language
YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit
pairs a short target utterance with the preceding context that makes its
figurative reading available, and carries a categorical label plus a free-text rationale.
This repository contains no audio. It ships annotations, transcriptions, and the
source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.tts-conversational-voice-20000h
TTS Voice Dataset
20,000 hours of high-fidelity 48kHz conversational audio across 30+ global, regional, and underrepresented languages, built for text-to-speech, voice cloning, and multilingual speech AI.
This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real audio samples.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/tts-conversational-voice-20000h.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.JobCCC-Conversational-Job-Recommendation-Bangladesh
JobCCC: A Conversational Code-Mixed Corpus for Job Recommendation in Bangladesh
Dataset Creators
Authors: Md. Arman Hossain, Mubashir Jawad, Fariha Khandakar Moon, and Sonia Binte Siraj
Supervisor: Dr. Nafis Sadeq
Institution: Department of Computer Science & Engineering, East West University
Dataset Summary
JobCCC (Conversational Code-Mixed Corpus) is a multi-turn conversational benchmark and job recommendation dataset tailored for the… See the full description on the dataset page: https://huggingface.co/datasets/Armans33115/JobCCC-Conversational-Job-Recommendation-Bangladesh.ndla_npk_conversational_nb_to_nn_hard
Hard dataset
This is a filtered hard version of the dataset NbAiLab/ndla_npk_conversational_nb_to_nn.
We have filtered out the 288793 of 1024909 lines where the models currently are making mistakes.
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.gemma3n-conversational-reasoning-with-tools
Gemma3N Conversational Reasoning With Embedded Tool Traces
Prepared for Unsloth Gemma3/Gemma3N conversational notebooks that expect ShareGPT conversations.
Multi-turn conversations are preserved.
Reasoning blocks (<think>...</think>) are preserved.
Tool call traces are preserved by embedding them in assistant text as tags:
<tool_call ...>...</tool_call>
<tool_response ...>...</tool_response>
Use:
from datasets import load_dataset
from unsloth.chat_templates import… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning-with-tools.prithivMLmods__QwQ-LCoT-14B-Conversational-details
Dataset Card for Evaluation run of prithivMLmods/QwQ-LCoT-14B-Conversational
Dataset automatically created during the evaluation run of model prithivMLmods/QwQ-LCoT-14B-Conversational
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__QwQ-LCoT-14B-Conversational-details.expresso-conversationalconversational-checkpoint-downloadsqrecc_conversational_embeddings
Dataset Card for "qrecc_conversational_embeddings"
More Information needed
gemma3n-conversational-reasoning-toolloop
Gemma3N Conversational Reasoning Tool-Loop
Gemma3N conversational dataset that preserves tool traces while avoiding training targets on tool responses.
Encoding:
Assistant emits tool calls: <tool_call ...>...</tool_call>
Tool outputs are user-side turns: <tool_response ...>...</tool_response>
This works with train_on_responses_only because user-side tool responses are masked from loss.
Use:
from datasets import load_dataset
from unsloth.chat_templates import… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning-toolloop.Expresso-Conversational-KanadeNemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/Mayur295/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.iskra-assistant-conversationalExpresso-Conversational-2Expresso-Conversational-Kanmirror-Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.vi-wiki-conversational-search
Dataset Card for Vi-Wiki-Conversational-Search
The ViWiki-QR dataset is a Vietnamese collection of 16.7K synthetic conversations and 250 human-annotated conversations, supporting the task of query rewriting for the conversational search.
Dataset Details
Dataset Description
ViWiki-QR is a Vietnamese dataset designed for the task of query rewriting in conversational search. It contains two subsets: a large-scale synthetic training set and a smaller, manually… See the full description on the dataset page: https://huggingface.co/datasets/trientp/vi-wiki-conversational-search.QuantumMechanics-6Chapters-ConversationalExpresso-Conversational-Processed-1Expresso-Conversational-Processed-2
