Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Roudranil /shakespearean-and-modern-english-conversational-dataset Dataset Card for Shakespearean and Modern English Conversational Dataset Dataset Summary This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details. text1K<n<10K4 likes224 downloads1y agoHugging Face02open-paws /conversational-finetuning-llama-format Open Paws Conversational Finetuning Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Training Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.texttext-generation10K<n<100K2 likes58 downloads1y agoHugging Face03Maral /conversational-persian-subtitles Conversational Persian Subtitles Dataset name: Conversational Persian SubtitlesCollaboration: Maral Zarvani & Milad Ghashangi AgdamLicense: CC BY 4.0Hugging Face Repo: https://huggingface.co/datasets/Maral/conversational-persian-subtitles 1. Dataset Description This dataset contains cleaned Persian subtitle lines from a wide variety of Korean TV series and films, each line reflecting informal, conversational dialogue. All markup (square brackets, timecodes,etc.) has… See the full description on the dataset page: https://huggingface.co/datasets/Maral/conversational-persian-subtitles.text100K<n<1M0 likes50 downloads1y agoHugging Face04darksyntax0 /conversational-sarcasm-benchmark Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit pairs a short target utterance with the preceding context that makes its figurative reading available, and carries a categorical label plus a free-text rationale. This repository contains no audio. It ships annotations, transcriptions, and the source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.tabularaudio-classification1K<n<10K0 likes50 downloads1mo agoHugging Face05jamesdborin /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.tabular10K<n<100K0 likes46 downloads3mo agoHugging Face06Ammad1Ali /Korean-conversational-datasettext10K<n<100K5 likes30 downloads3y agoHugging Face07Jeevan24 /Telugu-Conversational-Dataset-v1textn<1K1 likes23 downloads1y agoHugging Face08vetHealthGuy /Veteran_Affairs_NorthEast_Region_Conversational_Datasettext1K<n<10K0 likes17 downloads2y agoHugging Face09Bluestrikeai /Conversational-Hinglishtexttext-generationn<1K1 likes17 downloads1y agoHugging Face10ANASAKHTAR /bargaining_conversational_dataset.csvtext1K<n<10K0 likes15 downloads2y agoHugging Face11SoftAge-AI /sft-conversational_datasetgatedQuestion – Answer DatasetThe dataset contains 400 queries from two domains: Current Affairs and Creative Writing. It serves as a versatile resource for Natural Language Processing (NLP) tasks, including text classification, information retrieval, and model training. Data attributes: Query: The user-generated question. Data type: string. Answer: The response provided by a team of writers and editors in markdown format, containing information related to the query. Citations: Up to 4 credible… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/sft-conversational_dataset.textquestion-answeringn<1K5 likes14 downloads3y agoHugging Face12Zapd0s /Tanglish-Conversational-SFTA synthetic Tanglish (Tamil-English code-switched) conversational dataset for instruction tuning and supervised fine-tuning. text10K<n<100K0 likes14 downloads4mo agoHugging Face13AustinMcMike /steve_jobs_conversationaltextn<1K1 likes11 downloads3y agoHugging Face14propfirmkey /prop-trading-qa-conversational-ai Prop Trading Q&A Dataset for Conversational AI Description This dataset contains 200+ curated question-answer pairs covering the domain of proprietary (prop) trading firms. It is designed to serve as training and retrieval data for building AI assistants, chatbots, and educational tools focused on prop trading knowledge. Each entry consists of a natural-language question paired with a detailed, factual answer. The data spans ten thematic categories ranging from… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/prop-trading-qa-conversational-ai.textquestion-answeringn<1K0 likes10 downloads7mo agoHugging Face15avatrom /shakespearean-and-modern-english-conversational-dataset Dataset Card for Shakespearean and Modern English Conversational Dataset Dataset Summary This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details. text1K<n<10K0 likes10 downloads4mo agoHugging Face16trientp /vi-wiki-conversational-searchgated Dataset Card for Vi-Wiki-Conversational-Search The ViWiki-QR dataset is a Vietnamese collection of 16.7K synthetic conversations and 250 human-annotated conversations, supporting the task of query rewriting for the conversational search. Dataset Details Dataset Description ViWiki-QR is a Vietnamese dataset designed for the task of query rewriting in conversational search. It contains two subsets: a large-scale synthetic training set and a smaller, manually… See the full description on the dataset page: https://huggingface.co/datasets/trientp/vi-wiki-conversational-search.tabular100K<n<1M0 likes7 downloads1y agoHugging Face17bjw333 /ARGUS_Conversational_Datasettext1K<n<10K0 likes6 downloads2y agoHugging Face18longueirar /grammar_and_conversational_correctiontext1K<n<10K0 likes3 downloads2y agoHugging Face19pixend /en_fa_conversational_translationtext1M<n<10M1 likes3 downloads2y agoHugging Face20XXiaotao /conversationaltabular1K<n<10K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.