Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K48 likes199k downloads4d agoHugging Face02Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K8 likes111k downloads4d agoHugging Face03fka /prompts.chat a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit: 🌐 Website: prompts.chat 📦 GitHub: github.com/f/awesome-chatgpt-prompts About prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.textquestion-answering1K<n<10K9.9k likes27k downloads11h agoHugging Face04Goku-OpenLab /nano-banana-pro-prompts-datasets 🖼️ Nano Banana Pro Prompt Dataset 🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.imagetext-to-image10K<n<100K3 likes26k downloads2mo agoHugging Face05laion /voice-acting-cutscene-prompts Cut-Scene Voice-Acting Prompts Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting models. Each prompt describes a single speaker across two sharply contrasting emotional moments separated by a CUT TO: transition, in a voice-acting stage-direction format (spoken lines in "quotes", performance notes in (parentheses)). Total prompts: 4,057,000 Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.tabulartext-generation1M<n<10M2 likes17k downloads29d agoHugging Face06allenai /real-toxicity-prompts Dataset Card for Real Toxicity Prompts Dataset Summary RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models. Languages English Dataset Structure Data Instances Each instance represents a prompt and its metadata: { "filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt", "begin":340, "end":564, "challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.tabular10K<n<100K123 likes16k downloads4y agoHugging Face07HuggingFaceH4 /mt_bench_prompts MT Bench by LMSYS This set of evaluation prompts is created by the LMSYS org for better evaluation of chat models. For more information, see the paper. Dataset loading To load this dataset, use 🤗 datasets: from datasets import load_dataset data = load_dataset(HuggingFaceH4/mt_bench_prompts, split="train") Dataset creation To create the dataset, we do the following for our internal tooling. rename turns to prompts, add empty reference to remaining prompts… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/mt_bench_prompts.textquestion-answeringn<1K26 likes13k downloads3y agoHugging Face08codeShare /text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn This collection contains sets from the fusion-t2i-ai-generator on perchance. This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks To see the full sets, please use the url "https://perchance.org/" + url , where the urls are listed below: _generator gen_e621 fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.text-to-image100K<n<1M11 likes12k downloads2y agoHugging Face09Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes8.9k downloads2y agoHugging Face10stablellama /erotic-image-prompts Dataset Card for Erotic Image Prompts Dataset Description Dataset Summary Large language models (LLMs) are surprisingly bad at creatively inventing new things, even though they are masters of hallucination. Asking an LLM — even an abliterated one — to produce a list of random erotic prompts therefore yields a rather boring, narrow-minded result. This is easy to overcome: give the model some inspiration and let it do what it does best — transform… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/erotic-image-prompts.text-to-image10K<n<100K53 likes8.8k downloads18d agoHugging Face11TrustAIRLab /in-the-wild-jailbreak-prompts In-The-Wild Jailbreak Prompts on LLMs This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1,405… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts.tabulartext-generation10K<n<100K55 likes7.6k downloads2y agoHugging Face12artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K1 likes5.4k downloads2mo agoHugging Face13Gustavosta /Stable-Diffusion-Prompts Stable Diffusion Dataset This is a set of about 80,000 prompts filtered and extracted from the image finder for Stable Diffusion: "Lexica.art". It was a little difficult to extract the data, since the search engine still doesn't have a public API without being protected by cloudflare. If you want to test the model with a demo, you can go to: "spaces/Gustavosta/MagicPrompt-Stable-Diffusion". If you want to see the model, go to: "Gustavosta/MagicPrompt-Stable-Diffusion". text10K<n<100K532 likes4.9k downloads4y agoHugging Face14DeepPavlov /verbalist_prompts Verbalist (буквоед) - русскоязычный ассистент. Проект во многом вдохновленный Saiga. Мною были собраны все самые качественные датасеты с huggingface.datasets, а также собраны дополнительно с тех сайтов, которые я посчитал весьма полезными для создания аналога ChatGPT. Лицензии у всех датасетов отличаются, какие-то по типу OpenAssistant/oasst1 были созданы специально для обучения подобных моделей, какие-то являются прямой выгрузкой диалогов с ChatGPT (RyokoAI/ShareGPT52K). Вклад… See the full description on the dataset page: https://huggingface.co/datasets/DeepPavlov/verbalist_prompts.text1M<n<10M3 likes4.8k downloads3y agoHugging Face15OpenVoiceOS /ovos-tts-bench-massive-prompts OVOS tts bench — massive-prompts Synthesised clips (one per prompt) predictions of the registered OVOS Plugin Arena tts fighters over OpenVoiceOS/massive-templates. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-massive-prompts.0 likes3.7k downloads1mo agoHugging Face16AlekseyKorshuk /product-photography-v1-tiny-prompts-tasks-collage-filteredimage1K<n<10K1 likes2.6k downloads3y agoHugging Face17nateraw /parti-prompts Dataset Card for PartiPrompts (P2) Dataset Summary PartiPrompts (P2) is a rich set of over 1600 prompts in English that we release as part of this work. P2 can be used to measure model capabilities across various categories and challenge aspects. P2 prompts can be simple, allowing us to gauge the progress from scaling. They can also be complex, such as the following 67-word description we created for Vincent van Gogh’s The Starry Night (1889): Oil-on-canvas painting of a… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/parti-prompts.text1K<n<10K73 likes2.6k downloads4y agoHugging Face18rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K282 likes2.5k downloads3y agoHugging Face19erpgen /nsfw-writing-prompts1 likes2.4k downloads1y agoHugging Face20ChaoticNeutrals /Reddit-NSFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing [Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed. text1K<n<10K17 likes2.4k downloads2y agoHugging Face21Wenaka /Danbooru2024_rating_explicit_prompts_without_character精选了danbooru2024里分级为explict且score>20的所有图像的prompt,去除了原角色的tag和相关特征,可以直接用于角色nsfw图像生成 5 likes2.3k downloads1y agoHugging Face22XoxoMerlin /erotic-image-prompts Dataset Card for Erotic Image Prompts Dataset Description Dataset Summary Large language models (LLMs) are surprisingly bad at creatively inventing new things, even though they are masters of hallucination. Asking an LLM — even an abliterated one — to produce a list of random erotic prompts therefore yields a rather boring, narrow-minded result. This is easy to overcome: give the model some inspiration and let it do what it does best — transform… See the full description on the dataset page: https://huggingface.co/datasets/XoxoMerlin/erotic-image-prompts.text-to-image10K<n<100K0 likes2.3k downloads17d agoHugging Face23rx1lora /tb00-Reddit-NSFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing [Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed. 0 likes2.3k downloads2mo agoHugging Face24junvin /erotic-image-prompts Dataset Card for Erotic Image Prompts Dataset Description Dataset Summary Large language models (LLMs) are surprisingly bad at creatively inventing new things, even though they are masters of hallucination. Asking an LLM — even an abliterated one — to produce a list of random erotic prompts therefore yields a rather boring, narrow-minded result. This is easy to overcome: give the model some inspiration and let it do what it does best — transform… See the full description on the dataset page: https://huggingface.co/datasets/junvin/erotic-image-prompts.text-to-image10K<n<100K0 likes2.3k downloads18d agoHugging Face25lipilipic /Reddit-NSFW-Writing_Prompts_ShareGPT2 likes2.3k downloads6mo agoHugging Face26rx1lora /rev5-datas_Reddit-NSFW-Writing_Prompts_ShareGPT0 likes2.3k downloads2mo agoHugging Face27ganjaninja /writing-prompts-sfw-nsfw-interleaved0 likes2.3k downloads1y agoHugging Face28PromptSystematicReview /ThePromptReport Prompt Report Dataset This repository contains the dataset from the Prompt Report paper. Use huggingface hub or git lfs to download this, and use the instructions in our code repository to run the experiments. We also have a paper and website that detail our findings. master_papers.csv The master papers file is a master record of all the papers in the final dataset arxiv_papers_for_human_review.csv This csv contains the original group of papers… See the full description on the dataset page: https://huggingface.co/datasets/PromptSystematicReview/ThePromptReport.documentn<1K47 likes2.3k downloads2y agoHugging Face29Casual-Autopsy /Slop-Forensics_Nitral-AI_NSFW-SFW-Writing-Prompts-Mix-125x21 likes2.2k downloads2y agoHugging Face30jxcai-scale /hle_prompts_07_02_25text100K<n<1M0 likes2.2k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.