datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Finance-Conversational-Dataset-Indicmedical-symptom-triage-conversationalNemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.NCERT-Conversational-Dataset-IndicLaw-Conversational-Dataset-IndicCyber-Conversational-Dataset-Indicsmolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.conversational_dummyconversational_raw_datasetCoding-Conversational-Dataset-IndicComputer-Science-Conversational-Dataset-IndicCA-Conversational-Dataset-Indicemilia-mm-conversationalexpresso-conversational
The Expresso Dataset
[paper] [demo samples] [Original repository]
Introduction
The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided.
You can listen to… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/expresso-conversational.french-tts-conversational-dataset
French Conversational TTS Dataset
Dataset Description
This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution.
Verticals
Vertical
Description
fintech_banking
Banking operations, account inquiries, fraud alerts, investments, customer service
ecommerce_logistics
Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.japanese-casual-conversational-speech-golden-dataset-preview
Japanese Casual Conversational Speech Golden Dataset (Preview)
💼 Commercial License & Full Access
This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.
To purchase the full dataset, please contact us:
👉 Email: info@hth-inc.com
👉 Website: https://hth-inc.com/business
🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.conversational_weatherThe Conversational Weather dataset is designed for generation of responses to weather queries based on a structured input data. The input allows specifying data attributes such as dates, times, locations, weather conditions, and errors, and also offers control over structure of response through discourse relations such as join, contrast, and justification.NeMo-Gym-Conversational-Tool-Use-Assets
NeMo Gym Conversational Tool-Use Assets
This dataset repository stores prompt and reference assets for NeMo Gym's conversational tool-use generation pipeline.
It is an asset bundle for Gym components, not a training or evaluation dataset.
Contents
conversational_tool_use_domain_generation/prompts: the domain-generation prompt.
conversational_tool_use_domain_generation/prompt_history: historical domain-generation prompt revisions.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/NeMo-Gym-Conversational-Tool-Use-Assets.hle-no-img-conversational-formatTopical-Chat
Topical-Chat
We introduce Topical-Chat, a knowledge-grounded
human-human conversation dataset where the underlying
knowledge spans 8 broad topics and conversation
partners don’t have explicitly defined roles.
Topical-Chat broadly consists of two types of files:
Conversations: JSON files containing conversations between pairs of
Amazon Mechanical Turk workers.
Reading Sets: JSON files containing knowledge sections rendered as
reading content to the Turkers having conversations.
For… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-Chat.korean-conversational-speech-corpus
ASR-KCSC: A Korean Conversational Speech Corpus
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Speech Corpus
Language: Korean
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Recording Environment: Indoor
Dataset Description
This open-source dataset consists of 5.22 hours of transcribed Korean conversational speech on certain topics, where 22 conversations between seven pairs of speakers… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/korean-conversational-speech-corpus.Medical-Conversational-Dataset-Indicshakespearean-and-modern-english-conversational-dataset
Dataset Card for Shakespearean and Modern English Conversational Dataset
Dataset Summary
This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details.
conversational_data_untokenized_mergedconversational-s2s-v1
Conversational S2S v1 — dialogues parlés FR + EN
Corpus synthétique de conversations orales multi-tours entre un utilisateur
et un assistant, en français et en anglais, destiné au finetuning
speech-to-speech de LFM2.5-Audio (Liquid AI). Chaque tour est un clip audio
séparé, aligné avec son texte ; les dialogues sont conçus pour être rejoués tour
à tour (user → assistant → user → …).
Projet tts-model-exploration. Produit par
voxtral_datagen_pipeline (s2s-skeletons → remplissage… See the full description on the dataset page: https://huggingface.co/datasets/Rcarvalo/conversational-s2s-v1.Persian-conversational-datasetpersian-conversational-datasetConversationalRetrieval-SMAT
ReactiveAI / ConverstationalRetrieval Dataset for Supervised Memory-Aware Training (SMAT)
Description in progress
vietnamese-conversational-datasetconversational_audio_fr_dataset-metadataconversational-dynamics-egocom
Conversational Dynamics — EgoCom
Derived temporal annotations and model-ready training anchors for
conversational-dynamics and turn-taking research, generated from EgoCom.
This dataset is produced by the conversational-dynamics-data pipeline. The
underlying objective is to expose conversational data in a representation
suitable for temporal and action-conditioned models:
state_t + action_t → future conversational state
Contents
Three configurations are provided.… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom.
