Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stindardlogic /creative-writing-sft-50k Creative Writing SFT (50K) 50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure. Motivation Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.texttext-generation10K<n<100K0 likes7.1k downloads3mo agoHugging Face02Crownelius /Creative-Writing-Gemini3Pro-2700x Pulitzer Diamond Prose GEMINI Seeds This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.texttext-generation1K<n<10K5 likes4k downloads3mo agoHugging Face03Crownelius /Creative-Writing-Sonnet4.6-Cleaned Creative-Writing-Sonnet4.6-Cleaned Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.texttext-generationn<1K3 likes3.6k downloads3mo agoHugging Face04Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.9k downloads3mo agoHugging Face05BramVanroy /CommonCrawl-CreativeCommons The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.texttext-generation100M<n<1B42 likes2.5k downloads1y agoHugging Face06yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.3k downloads7mo agoHugging Face07Aratako /Japanese-Creative-Writing-39.6k Japanese-Creative-Writing-39.6k 概要 deepseek-ai/DeepSeek-V3-0324を用いて作成した、約39600件の日本語の小説執筆タスクデータセットです。 全てのデータは2ターンのデータとなっています。また、データセット中の一部データはNSFW表現を含みます。 データの詳細 各データは以下のキーを含みます。 messages: OpenAI messages形式の対話データ instruction_1: 1ターン目の指示プロンプト output_1: 1ターン目のアシスタント応答 instruction_2: 2ターン目の指示プロンプト output_2: 2ターン目のアシスタント応答 1ターン目の指示プロンプトはdeepseek-ai/DeepSeek-V3-0324で合成されています。system promptや2ターン目の指示プロンプトは事前に用意した複数種類からランダムに選択されたものが設定されています。 ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-Creative-Writing-39.6k.texttext-generation10K<n<100K8 likes1.1k downloads1y agoHugging Face08Crownelius /Creative-Writing-Qwen3.5Plus-2000x Pulitzer Diamond Prose QWEN Seeds This dataset contains 2638 high-quality creative writing seeds generated using Qwen 2.5 72B. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Qwen3.5Plus-2000x.texttext-generation1K<n<10K2 likes922 downloads3mo agoHugging Face09BramVanroy /CommonCrawl-CreativeCommons-fine Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46 CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.texttext-generation10M<n<100M5 likes911 downloads1y agoHugging Face10Crownelius /Creative-Writing-Sonnet4.6-800x Pulitzer Diamond Prose CLAUDE Seeds This dataset contains 833 high-quality creative writing seeds generated using Claude 4.6 Sonnet. Each entry represents a story opening designed to meet high literary standards. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements: extreme show-don't-tell, double-labor sentence… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-800x.texttext-generationn<1K7 likes900 downloads3mo agoHugging Face11Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes838 downloads3mo agoHugging Face12BramVanroy /CommonCrawl-CreativeCommons-strict Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.texttext-generation10M<n<100M2 likes794 downloads1y agoHugging Face13Crownelius /Creative-Writing-Part-Two Creative Writing - Part Two (The Nuclear Dataset) This dataset represents the "Nuclear" layer of our creative writing training pipeline. While Part One focused on physical and psychological grounding (Shadow & Skeleton), Part Two focuses on dense literary resonance, subtext, and stylistic sophistication. Methodology: The Nuclear Pipeline This dataset was built using a multi-phase "Controlled Criticality" approach to ensure maximum signal density without the… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Part-Two.texttext-generation1K<n<10K3 likes674 downloads3mo agoHugging Face14rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes554 downloads7mo agoHugging Face15kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes474 downloads7mo agoHugging Face16Crownelius /Creative_Writing_ShareGPT_Enhanced Creative Writing ShareGPT — Enhanced Edition ✨ High-quality creative writing dataset with regenerated responses using StepFun's Step-3.5-Flash model. This dataset is an enhanced version of ChaoticNeutrals/Creative_Writing-ShareGPT, where all final AI responses have been regenerated using stepfun/step-3.5-flash with a carefully engineered system prompt designed to produce literary-quality creative writing. What Changed Original human prompts preserved — All… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative_Writing_ShareGPT_Enhanced.text-generation1K<n<10K5 likes270 downloads3mo agoHugging Face17BreadStudio /cqa-creative-writing-expert-cot-preview CQA: Creative Quality Alignment — Research-Grade Schema v2 English This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.texttext-generationn<1K6 likes234 downloads3mo agoHugging Face18Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes153 downloads3mo agoHugging Face19telecomadm1145 /creative_writing Dataset Card for telecomadm1145/creative_writing Dataset Details Dataset Description This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output). The dataset emphasizes: Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.texttext-generation1K<n<10K6 likes130 downloads1y agoHugging Face20nchapman /figaro-creative-writing figaro-creative-writing A high-quality creative writing dataset built using an editor feedback pipeline. Each story goes through three stages: DeepSeek V3.2 writes a first draft, Grok 4.1 Fast provides detailed editorial feedback, then DeepSeek revises based on that feedback. The revision is the final output. Overview Rows 6,022 Format Single-turn chat (system + user + assistant) Prompts Gryphe/Opus-WritingPrompts Writer model DeepSeek V3.2 (temp=1.0… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/figaro-creative-writing.texttext-generation1K<n<10K0 likes118 downloads7mo agoHugging Face21Crownelius /Creative_Writing_Multiturn_Enhanced Creative Writing Multiturn — Enhanced Edition ✨ High-quality creative writing dataset with regenerated responses using StepFun's Step-3.5-Flash model. This dataset is an enhanced version of Dampfinchen/Creative_Writing_Multiturn, where all final AI responses have been regenerated using stepfun/step-3.5-flash with a carefully engineered system prompt designed to produce literary-quality creative writing. What Changed Original human prompts preserved — All user… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative_Writing_Multiturn_Enhanced.text-generation1K<n<10K3 likes114 downloads3mo agoHugging Face22DarkyMan /Opus-4.6-RU-Reasoning-creative-1385x-not-filtered Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response. Dataset Info Language: Russian 🇷🇺 Size: ~1,385 samples (growing) Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"} Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.texttext-generation1K<n<10K4 likes108 downloads7mo agoHugging Face23vicgalle /creative-rubrics creative-rubrics 🎏 A dataset of creative responses using GPT-4.5, o3-mini and DeepSeek-R1. This dataset contains several prompts seeking creative and diverse answers (like writing movie reviews, short stories, etc), and the style of the responses has been enhanced by prompting the model with custom rubrics that seek different creative styles. It can be used for finetuning for custom styles with open-text tasks. The dataset was presented in the paper Configurable Preference Tuning… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/creative-rubrics.texttext-generation1K<n<10K8 likes104 downloads1y agoHugging Face24enPurified /smoltalk-creative-writing-enPurified-openai-messages 📖 SmolTalk-Creative-Writing-enPurified-openai-messages SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset. The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.texttext-generation10K<n<100K2 likes94 downloads9mo agoHugging Face25SolusOps /incremental-instruction-creative-writinggated Incremental Instruction Creative Writing Does delivering a writing brief over several conversation turns change what a language model writes? This dataset supports that question with matched creative-writing tasks evaluated under two delivery conditions: FULL: the complete brief is supplied in one turn. SHARDED: the same intended brief is introduced across five to nine turns. The benchmark holds task content fixed while varying how the instructions are delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.tabulartext-generation1K<n<10K0 likes92 downloads1mo agoHugging Face26AngelWarmSmile123 /deep-creative-writing-zh Deep Creative Writing Dialogue Dataset (Chinese) 深度文学创作对话数据集 Dataset Description High-quality Chinese creative writing dialogues covering novel structure, character development, narrative techniques, symbolism, and literary theory. 高质量中文文学创作对话,涵盖小说结构设计、角色塑造、叙事技巧、象征主义、文学理论等议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-creative-writing-zh.texttext-generation1K<n<10K1 likes85 downloads3mo agoHugging Face27vicgalle /creative-rubrics-preferences creative-rubrics-preferences 🎏 A dataset of creative responses using GPT-4.5, o3-mini and DeepSeek-R1. This dataset contains several prompts seeking creative and diverse answers (like writing movie reviews, short stories, etc), and the style of the responses has been enhanced by prompting the model with custom rubrics that seek different creative styles. This dataset was used in the paper Configurable Preference Tuning with Rubric-Guided Synthetic Data. Code:… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/creative-rubrics-preferences.texttext-generationn<1K4 likes83 downloads1y agoHugging Face28oliveirabruno01 /ptbr-creative-cpt-qwen35-08b-v02 PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.tabulartext-generation1K<n<10K0 likes83 downloads19d agoHugging Face29oliveirabruno01 /ptbr-creative-cpt PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.tabulartext-generation1K<n<10K0 likes77 downloads19d agoHugging Face30ZachW /gemma-4-31b-it_arena-hard-creative-writing google/gemma-4-31b-it — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_arena-hard-creative-writing.tabulartext-generationn<1K1 likes76 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.