datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shakespearean-and-modern-english-conversational-dataset
Dataset Card for Shakespearean and Modern English Conversational Dataset
Dataset Summary
This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details.
conversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.conversational-persian-subtitles
Conversational Persian Subtitles
Dataset name: Conversational Persian SubtitlesCollaboration: Maral Zarvani & Milad Ghashangi AgdamLicense: CC BY 4.0Hugging Face Repo: https://huggingface.co/datasets/Maral/conversational-persian-subtitles
1. Dataset Description
This dataset contains cleaned Persian subtitle lines from a wide variety of Korean TV series and films, each line reflecting informal, conversational dialogue. All markup (square brackets, timecodes,etc.) has… See the full description on the dataset page: https://huggingface.co/datasets/Maral/conversational-persian-subtitles.conversational-sarcasm-benchmark
Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only
A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language
YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit
pairs a short target utterance with the preceding context that makes its
figurative reading available, and carries a categorical label plus a free-text rationale.
This repository contains no audio. It ships annotations, transcriptions, and the
source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.Korean-conversational-datasetTelugu-Conversational-Dataset-v1Veteran_Affairs_NorthEast_Region_Conversational_DatasetConversational-Hinglishbargaining_conversational_dataset.csvsft-conversational_datasetQuestion – Answer DatasetThe dataset contains 400 queries from two domains: Current Affairs and Creative Writing. It serves as a versatile resource for Natural Language Processing (NLP) tasks, including text classification, information retrieval, and model training.
Data attributes:
Query: The user-generated question. Data type: string.
Answer: The response provided by a team of writers and editors in markdown format, containing information related to the query.
Citations: Up to 4 credible… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/sft-conversational_dataset.Tanglish-Conversational-SFTA synthetic Tanglish (Tamil-English code-switched) conversational dataset for instruction tuning and supervised fine-tuning.
steve_jobs_conversationalprop-trading-qa-conversational-ai
Prop Trading Q&A Dataset for Conversational AI
Description
This dataset contains 200+ curated question-answer pairs covering the domain of proprietary (prop) trading firms. It is designed to serve as training and retrieval data for building AI assistants, chatbots, and educational tools focused on prop trading knowledge.
Each entry consists of a natural-language question paired with a detailed, factual answer. The data spans ten thematic categories ranging from… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/prop-trading-qa-conversational-ai.shakespearean-and-modern-english-conversational-dataset
Dataset Card for Shakespearean and Modern English Conversational Dataset
Dataset Summary
This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details.
vi-wiki-conversational-search
Dataset Card for Vi-Wiki-Conversational-Search
The ViWiki-QR dataset is a Vietnamese collection of 16.7K synthetic conversations and 250 human-annotated conversations, supporting the task of query rewriting for the conversational search.
Dataset Details
Dataset Description
ViWiki-QR is a Vietnamese dataset designed for the task of query rewriting in conversational search. It contains two subsets: a large-scale synthetic training set and a smaller, manually… See the full description on the dataset page: https://huggingface.co/datasets/trientp/vi-wiki-conversational-search.ARGUS_Conversational_Datasetgrammar_and_conversational_correctionen_fa_conversational_translationconversational
