Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sweatSmile /medical-symptom-triage-conversationaltabular1K<n<10K1 likes1.6k downloads1y agoHugging Face02nvidia /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K42 likes1.5k downloads12d agoHugging Face03oss-codes /Finance-Conversational-Dataset-Indictext100K<n<1M1 likes913 downloads2y agoHugging Face04ElementXMaster /conversational_raw_datasettext10M<n<100M0 likes900 downloads1y agoHugging Face05oss-codes /Law-Conversational-Dataset-Indictext100K<n<1M0 likes690 downloads2y agoHugging Face06AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes677 downloads2mo agoHugging Face07oss-codes /NCERT-Conversational-Dataset-Indictext100K<n<1M0 likes611 downloads2y agoHugging Face08batgre /conversational-dynamics-egocom Conversational Dynamics — EgoCom Derived temporal annotations and model-ready training anchors for conversational-dynamics and turn-taking research, generated from EgoCom. This dataset is produced by the conversational-dynamics-data pipeline. The underlying objective is to expose conversational data in a representation suitable for temporal and action-conditioned models: state_t + action_t → future conversational state Contents Three configurations are provided.… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom.tabular1M<n<10M1 likes521 downloads5d agoHugging Face09soda-research /emilia-mm-conversationalgatedtext100K<n<1M0 likes488 downloads10mo agoHugging Face10nytopop /expresso-conversational The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided. You can listen to… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/expresso-conversational.audio10K<n<100K14 likes465 downloads1y agoHugging Face11HTH-inc /japanese-casual-conversational-speech-golden-dataset-preview Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.audioautomatic-speech-recognitionn<1K2 likes417 downloads1mo agoHugging Face12Rcarvalo /conversational-s2s-v1 Conversational S2S v1 — dialogues parlés FR + EN Corpus synthétique de conversations orales multi-tours entre un utilisateur et un assistant, en français et en anglais, destiné au finetuning speech-to-speech de LFM2.5-Audio (Liquid AI). Chaque tour est un clip audio séparé, aligné avec son texte ; les dialogues sont conçus pour être rejoués tour à tour (user → assistant → user → …). Projet tts-model-exploration. Produit par voxtral_datagen_pipeline (s2s-skeletons → remplissage… See the full description on the dataset page: https://huggingface.co/datasets/Rcarvalo/conversational-s2s-v1.audioaudio-to-audio10K<n<100K0 likes408 downloads1mo agoHugging Face13oss-codes /Cyber-Conversational-Dataset-Indictext1K<n<10K0 likes401 downloads2y agoHugging Face14ksterx /hle-no-img-conversational-formatimage1K<n<10K0 likes366 downloads1y agoHugging Face15nvidia /NeMo-Gym-Conversational-Tool-Use-Assets NeMo Gym Conversational Tool-Use Assets This dataset repository stores prompt and reference assets for NeMo Gym's conversational tool-use generation pipeline. It is an asset bundle for Gym components, not a training or evaluation dataset. Contents conversational_tool_use_domain_generation/prompts: the domain-generation prompt. conversational_tool_use_domain_generation/prompt_history: historical domain-generation prompt revisions.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/NeMo-Gym-Conversational-Tool-Use-Assets.textn<1K2 likes366 downloads2mo agoHugging Face16batgre /conversational-dynamics-egocom-12.5hz Conversational Dynamics — EgoCom (12.5 Hz) Derived temporal annotations and model-ready training anchors for conversational-dynamics and turn-taking research, generated from EgoCom. This dataset is produced by the conversational-dynamics-data pipeline. The underlying objective is to expose conversational data in a representation suitable for temporal and action-conditioned models: state_t + action_t → future conversational state 12.5 Hz release The same pipeline… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom-12.5hz.tabular1M<n<10M0 likes339 downloads5d agoHugging Face17oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes294 downloads2y agoHugging Face18MagicHub /korean-conversational-speech-corpus ASR-KCSC: A Korean Conversational Speech Corpus Every data point counts. Dataset Basic Info Dataset Type: ASR Speech Corpus Language: Korean Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Recording Environment: Indoor Dataset Description This open-source dataset consists of 5.22 hours of transcribed Korean conversational speech on certain topics, where 22 conversations between seven pairs of speakers… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/korean-conversational-speech-corpus.audio1K<n<10K1 likes270 downloads4mo agoHugging Face19oss-codes /CA-Conversational-Dataset-Indictext100K<n<1M0 likes266 downloads2y agoHugging Face20oss-codes /Medical-Conversational-Dataset-Indictext10K<n<100K0 likes250 downloads2y agoHugging Face21Roudranil /shakespearean-and-modern-english-conversational-dataset Dataset Card for Shakespearean and Modern English Conversational Dataset Dataset Summary This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details. text1K<n<10K4 likes224 downloads1y agoHugging Face22ElementXMaster /conversational_data_untokenized_mergedtext1M<n<10M0 likes206 downloads1y agoHugging Face23Kamtera /Persian-conversational-datasetpersian-conversational-datasettexttext-generation100K<n<1M12 likes191 downloads4y agoHugging Face24ReactiveAI /ConversationalRetrieval-SMAT ReactiveAI / ConverstationalRetrieval Dataset for Supervised Memory-Aware Training (SMAT) Description in progress text100K<n<1M0 likes184 downloads5mo agoHugging Face25Makan09 /bam-asr-conversational All Bambara ASR Dataset This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-conversational.audioautomatic-speech-recognition10K<n<100K2 likes164 downloads1mo agoHugging Face26malaysia-ai /malay-conversational-speech-corpus malay-conversational-speech-corpus Mirror for https://magichub.com/datasets/malay-conversational-speech-corpus/, license is Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License audio1K<n<10K5 likes141 downloads3y agoHugging Face27B-Lounes /conversational_audio_fr_dataset-metadatatextn<1K0 likes136 downloads1y agoHugging Face28mpasila /Gemma-3-Finnish-ShareGPT-Conversational-Dataset-V1The "topics" folder contains each topic generated (each was supposed to have 5 different conversations but some may have failed to generate or were removed due to big mistakes in them). Some stuff may contain NSFW content. The main dataset file has all topics shuffled into one .json file. Scripts used when creating the dataset: https://github.com/minipasila/dataset-scripts/tree/main/gemma-3-dataset-stuff There have been some small manual edits but it's mostly unedited. This is more of a… See the full description on the dataset page: https://huggingface.co/datasets/mpasila/Gemma-3-Finnish-ShareGPT-Conversational-Dataset-V1.textn<1K0 likes132 downloads1y agoHugging Face29Aobangaming /Conversational-Fine-Tuning Dataset Card for CFT Conversational Fine-tuning is a dataset meant for fine-tuning conversational models. This dataset contains prompts generated by ChatGPT and Gemini. Dataset Details Dataset Description Conversational Fine-Tuning aims to be a basic fine-tuning dataset for large language models in early stages of conversational training. The dataset has over 5000+ unique tokens generated by AI language models. Curated by: AobanZ Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Aobangaming/Conversational-Fine-Tuning.text1K<n<10K3 likes130 downloads28d agoHugging Face30harryxi /reward-bench-by-subset-conversationaltext1K<n<10K0 likes120 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.