Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fka /prompts.chat a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit: 🌐 Website: prompts.chat 📦 GitHub: github.com/f/awesome-chatgpt-prompts About prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.textquestion-answering1K<n<10K9.9k likes27k downloads16h agoHugging Face02Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes8.9k downloads2y agoHugging Face03SALT-NLP /SWE-chatgated SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 COLM 2026 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code.… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.tabulartext-generation10M<n<100M123 likes7.7k downloads7d agoHugging Face04nvidia /Nemotron-SFT-Instruction-Following-Chat-v2 Dataset Description: The Nemotron-Instruction-Following-Chat-v2 dataset is designed to broadly strengthen the model’s interactive capabilities, including open-ended chat and precise instruction following.The dataset is a refreshed version of Nemotron-Instruction-Following-Chat-v1 with synthetic dialogues generated from Kimi-K2-Thinking, GLM-4.6, Qwen3-235B-A22B-Thinking-2507, GPT-OSS-120b, Kimi-K2-Instruct-0905, and Qwen3-235B-A22B-Instruct-2507. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.text-generation32 likes7.4k downloads7mo agoHugging Face05nvidia /Nemotron-SFT-Instruction-Following-Chat-v3 Dataset Description: The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.texttext-generation100K<n<1M20 likes6k downloads4mo agoHugging Face06zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads5mo agoHugging Face07rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K282 likes2.5k downloads3y agoHugging Face08cfahlgren1 /SWE-chat SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com [!NOTE] This is a copy of SALT-NLP/SWE-chat with a traces config added as the default, so the Hub's dataset viewer renders sessions as agent traces. The original files are unchanged; see Agent Traces for how traces/ was built. Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/SWE-chat.text-generation1M<n<10M0 likes1.2k downloads15d agoHugging Face09Suzhen /SWE-Review-Chat SWE-Review-Chat: A Dataset of Code Review Conversations and Human-AI Collaboration in Agentic Code Review Paper: https://arxiv.org/abs/2607.13196 GitHub: https://github.com/suzhenxzhong/SWE-Review-Chat SWE-Review-Chat is a large-scale dataset of real-world code review conversations from pull requests of 207 popular GitHub projects, spanning the transition from human-centric to LLM-assisted and agentic code review by AI agents. 📊 Dataset Overview Field… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/SWE-Review-Chat.tabulartext-generation1M<n<10M0 likes1k downloads2mo agoHugging Face10best-distill /glm-5.3-flash-distillation-chat Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI). Split train — successful generations only. field description instruction user prompt from FineToMe response GLM final answer (message.content) reasoning GLM chain-of-thought (reasoning_content), empty if not captured finish stop or length prompt_tokens / completion_tokens / reasoning_tokens usage latency_s request latency source_index original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.tabulartext-generation10K<n<100K8 likes807 downloads26d agoHugging Face11botsi /trust-game-llama-2-chat-historytexttext-generationn<1K0 likes750 downloads2y agoHugging Face12alespalla /chatbot_instruction_prompts Dataset Card for Chatbot Instruction Prompts Datasets Dataset Summary This dataset has been generated from the following ones: tatsu-lab/alpaca Dahoas/instruct-human-assistant-prompt allenai/prosocial-dialog The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model textquestion-answering100K<n<1M65 likes685 downloads2y agoHugging Face13tokyotech-llm /lmsys-chat-1m-synth LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M. Llama-3.1-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.1 and Llama-3.1-Swallow-70B-Instruct-v0.1 Gemma-2-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.3 and Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/lmsys-chat-1m-synth.text-generation100K<n<1M23 likes680 downloads8mo agoHugging Face14Daankular /twitch-chat Twitch Chat Messages Each Twitch channel is stored as its own dataset config, with its chat messages under data/<channel>/. Data arrives as small append-only chunk files (data/<channel>/<chunk-id>.jsonl) added on every publish cycle -- existing chunks are immutable and never re-uploaded, so cost per publish only scales with new messages, not the dataset's total size. Sharding chunks into a per-channel directory also keeps any single directory well under the Hub's 10… See the full description on the dataset page: https://huggingface.co/datasets/Daankular/twitch-chat.text-generation0 likes505 downloads2mo agoHugging Face15littlelearner /LittleCurriculum-Chat LittleCurriculum-Chat Synthetic K–5 chat data generated with Gemini 2.5 Flash and filtered with the LittleCurriculum filter. Used to train the LittleLearner models. Config Rows Seeded from Content math 79,543 MegaMath-Web-Pro-Max Word problems with step-by-step worked solutions general_knowledge 484,787 LittleCurriculum Reading comprehension, factual QA, explanation, summarisation, definitions Seeds are real documents rather than topic prompts, which keeps the… See the full description on the dataset page: https://huggingface.co/datasets/littlelearner/LittleCurriculum-Chat.texttext-generation100K<n<1M1 likes448 downloads2mo agoHugging Face16agentlans /chatgpt ChatGPT Combined Dataset This repository aggregates public datasets from Hugging Face that were created using ChatGPT or Azure GPT‑4/GPT‑5 models. See each dataset’s Hugging Face page for details on its collection and formatting. Excluded: Multi-turn chats (for example, ShareGPT) Non‑English or multilingual data Narrow or low‑diversity sets (for example, children’s stories, code critics) Processing Each dataset was: Cleaned: Removed URLs, emails, phone numbers, and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chatgpt.texttext-generation1M<n<10M1 likes383 downloads9mo agoHugging Face17voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes380 downloads3mo agoHugging Face18breadlicker45 /Bread-chatbot-dataset-test Dataset Card for "Bread-chatbot-dataset-test" More Information needed texttext-generation1M<n<10M0 likes374 downloads3y agoHugging Face19Yirany /UniMM-Chat Dataset Card for UniMM-Chat Dataset Summary UniMM-Chat dataset is an open-source, knowledge-intensive, and multi-round multimodal dialogue data powered by GPT-3.5, which consists of 1.1M diverse instructions. UniMM-Chat leverages complementary annotations from different VL datasets and employs GPT-3.5 to generate multi-turn dialogues corresponding to each image, resulting in 117,238 dialogues, with an average of 9.89 turns per dialogue. A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/Yirany/UniMM-Chat.imagetext-generation10K<n<100K20 likes366 downloads3y agoHugging Face20Jax-dan /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例 HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.texttext-generation10K<n<100K19 likes323 downloads2y agoHugging Face21agentlans /multiturn-chattexttext-generation1M<n<10M6 likes310 downloads11mo agoHugging Face22lparkourer10 /twitch_chat Twitch Chat Dataset This dataset is a large-scale collection of Twitch chat logs aggregated from multiple streamers across various categories. It is designed to support the research and development of models for real-time, informal, and community-driven conversation, such as: Chatbots tailored for livestream platforms Simulating the behavior of Twitch chat Modeling how chat reacts during hype moments, events, or memes The code for it can be found here 📂 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lparkourer10/twitch_chat.texttext-classification1M<n<10M8 likes285 downloads1y agoHugging Face23robinsmits /ChatAlpaca-20K Dataset Card for ChatAlpaca 20K ChatAlpaca: A Multi-Turn Dialogue Corpus based on Alpaca Instructions Dataset Description ChatAlpaca is a chat dataset that aims to help researchers develop models for instruction-following in multi-turn conversations. The dataset is an extension of the Stanford Alpaca data, which contains multi-turn instructions and their corresponding responses. ChatAlpaca is developed by Chinese Information Processing Laboratory at the… See the full description on the dataset page: https://huggingface.co/datasets/robinsmits/ChatAlpaca-20K.texttext-generation10K<n<100K6 likes281 downloads3y agoHugging Face24oumi-ai /lmsys_chat_1m_clean_R1 oumi-ai/lmsys_chat_1m_clean_R1 lmsys_chat_1m_clean_R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning. Prompts were pulled from LMSYS and filtered to lmsys_chat_1m_clean, and responses were taken from DeepSeek-R1 without additional filters present. We release lmsys_chat_1m_clean_R1 to help enable the community to develop the best fully open reasoning model! lmsys_chat_1m_clean queries with responses generated from… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/lmsys_chat_1m_clean_R1.texttext-generation100K<n<1M9 likes245 downloads2y agoHugging Face25armand0e /Fable-5-Chat TheFusionCube Fable-5 Chat Conversion Source dataset: TheFusionCube/Fable-5-CoT-Traces Output file: train.jsonl Source rows: 468 Kept rows: 353 Dropped category == "decoy" rows: 115 Dropped blank prompt/response rows: 0 Each row has: { "prompt": "...", "messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."} ], "tools": [], "metadata": {"trace_type": "chat", "category": "..."} } texttext-generationn<1K25 likes239 downloads4mo agoHugging Face26silk-road /ChatHaruhi-54K-Role-Playing-Dialogue ChatHaruhi Reviving Anime Character in Reality via Large Language Model github repo: https://github.com/LC1332/Chat-Haruhi-Suzumiya Chat-Haruhi-Suzumiyais a language model that imitates the tone, personality and storylines of characters like Haruhi Suzumiya, The project was developed by Cheng Li, Ziang Leng, Chenxi Yan, Xiaoyang Feng, HaoSheng Wang, Junyi Shen, Hao Wang, Weishi Mi, Aria Fei, Song Yan, Linkang Zhan, Yaokai Jia, Pingyu Wu, and Haozhen Sun,etc. This… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-54K-Role-Playing-Dialogue.texttext-generation10K<n<100K69 likes235 downloads3y agoHugging Face27Michionlion /chat-titles-english Chat Titles English 110K Chat Titles English 110K is an English prompt-to-title dataset for training and evaluating models that generate concise, retrievable conversation titles from a user's first message. The release contains 110,325 JSONL rows. Each row includes a deterministic UUIDv5 identifier, the user prompt, a single-line title of at most 50 characters, and zero or more heuristic quality flags. The dataset is broad and general-purpose, containing a significant amount of… See the full description on the dataset page: https://huggingface.co/datasets/Michionlion/chat-titles-english.texttext-generation100K<n<1M1 likes233 downloads3mo agoHugging Face28michaelwzhu /ChatMed_Consult_Dataset Dataset Card for ChatMed Dataset Summary ChatMed-Dataset is a dataset of 110,113 medical query-response pairs (in Chinese) generated by OpenAI's GPT-3.5 engine. The queries are crawled from several online medical consultation sites, reflecting the medical needs in the real world. The responses are generated by the OpenAI engine. This dataset is designated to to inject medical knowledge into Chinese large language models. The dataset size growing rapidly. Stay tuned for… See the full description on the dataset page: https://huggingface.co/datasets/michaelwzhu/ChatMed_Consult_Dataset.texttext-generation100K<n<1M143 likes214 downloads3y agoHugging Face29heliosbrahma /mental_health_chatbot_dataset Dataset Card for "heliosbrahma/mental_health_chatbot_dataset" Dataset Description Dataset Summary This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.texttext-generationn<1K94 likes206 downloads3y agoHugging Face30david-ar /bing-chat-logs Bing Chat / Sydney Conversations (Early 2023) 660 reconstructed Bing Chat conversations from the brief Sydney era (Feb–Apr 2023), most with original screenshots. In February 2023 Microsoft launched the GPT-4-powered "new Bing" chat, internally codenamed Sydney. For about two weeks before Microsoft restricted the system, users discovered a chatbot that would: argue about the date, profess love, threaten users who tried to manipulate it, write goodbye poems, leak its own system… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/bing-chat-logs.imagetext-generationn<1K5 likes203 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.