Team Ai
Datasetpublic

Travis-ML/ShortStory-SFT-jsonl

Public Domain Short Fiction with Prompts 719 complete short stories (300 to 2500 words) by 26 authors whose work is in the public domain, each paired with a natural-language request that could plausibly have produced it. Built for supervised fine-tuning of small language models on fiction, where the usual sources (forum stories, model-generated stories) lack the structural control of published short fiction. Fields field description id stable id (hash… See the full description on the dataset page: https://huggingface.co/datasets/Travis-ML/ShortStory-SFT-jsonl.

sourceHugging Facecc0-1.0updated 20d agoView on Hugging Face
0likes101downloads
Dataset Card

Public Domain Short Fiction with Prompts

719 complete short stories (300 to 2500 words) by 26 authors whose work is in the public domain, each paired with a natural-language request that could plausibly have produced it. Built for supervised fine-tuning of small language models on fiction, where the usual sources (forum stories, model-generated stories) lack the structural control of published short fiction.

Fields

fielddescription
idstable id (hash of author + title)
authorauthor as catalogued by the source
bookthe collection or volume the story was taken from
titlestory title
wordsword count of text
prompt_briefa short request (1-2 sentences): premise and setup only, sometimes a length target; never the author or title
prompt_detaileda fuller request (2-4 sentences): premise, tone, point of view, setting or era; sometimes a length target; often outlines the arc (see caveats)
genresup to 4 genre or mood tags
kindstory: every text was classified by Claude as a complete, self-contained short story (chapters, essays and verse were dropped)
textthe story, one paragraph per line, blank line between paragraphs
licenseprovenance note for the text

For chat-format SFT, use either prompt (or sample between them) as the user turn and text as the assistant turn.

How it was built

Texts come from public-domain plain-text editions of short story collections (transcribed by Project Gutenberg volunteers; see https://www.gutenberg.org). Each collection was split into stories using its table of contents, transcription headers/footers and illustration markers were removed, hard line wraps were undone, and novels (numbered chapters), verse and non-fiction were excluded. Stories appearing in more than one collection were kept once.

Each text was then classified by Claude (claude-haiku-4-5) as a complete short story, a novel chapter, an essay, a poem, or other; only stories are included. Prompts were written by the same model shown the full story and asked for the request a person might have typed to get it, in two styles (brief and detailed), without naming the author, title or source, and without revealing the ending. About half the prompts carry a length target matching the story. Prompts were spot-checked, not fully hand-reviewed. Two known weaknesses: the detailed prompts frequently outline the whole arc (sometimes including the ending) despite being asked not to, so they read more like a synopsis than a casual request; and phrasing is less varied than intended ("Write a ..." opens about half the brief prompts and most of the detailed ones). If you want prompt diversity, mix with other sources or rewrite the prompts.

Composition

Length (words): 1000-1499: 173, 1500+: 438, 300-499: 8, 500-999: 100

Top authors: Lang (150), Andersen (98), Chekhov (98), Grimm (89), Maupassant (82), Kipling (41), Baum (33), Nesbit (20), Poe (16), Twain (16), Potter (10), Hawthorne (9), Wells (8), London (7), Doyle (6)

Caveats

Most texts are from the 1830s to 1920s and carry the register of their period; a model trained only on this set will write like it. Mix with contemporary data if that is not what you want. Some stories reflect attitudes of their time. Fairy tales and children's stories (Lang's colored fairy books, Andersen, Grimm, Baum, Nesbit, Potter) are over half the set: they are what survives a filter for complete stories under 2,500 words from this period. The rest is mostly Chekhov, Maupassant, Kipling, Poe, Twain and Hawthorne.

License

Texts: public domain. Prompts and tags: released under CC0 1.0. The transcriptions were obtained from Project Gutenberg; this dataset is not affiliated with or endorsed by Project Gutenberg.