Inkwell-Software/screenplay-format-edge-cases
Screenplay Format Edge Cases 48 original Fountain specimens in 24 contrast pairs — two near-identical inputs per pair, at the points where the Fountain syntax leaves a choice. In 17 pairs the one difference changes how the lines are classified. In the other 7 it changes the surface and the labels hold: a lowercase scene prefix, a cue extension, a non-Latin cue, escaped characters, a dual-dialogue caret, an inline note and centered-text markers. Version: 1.0.0 · Maintainer:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-format-edge-cases.
Screenplay Format Edge Cases
48 original Fountain specimens in 24 contrast pairs — two near-identical inputs per pair, at the points where the Fountain syntax leaves a choice. In 17 pairs the one difference changes how the lines are classified. In the other 7 it changes the surface and the labels hold: a lowercase scene prefix, a cue extension, a non-Latin cue, escaped characters, a dual-dialogue caret, an inline note and centered-text markers.
Version: 1.0.0 · Maintainer: Inkwell.
Built by Inkwell, the IDE for screenwriters. Paste Fountain into the Fountain validator to see the scenes, character cues, dialogue and action it recognizes.
What's included
data/cases.jsonl holds all 48 records; schema.json is a JSON Schema for one record. Each record carries the Fountain input, its expected line-level element types, the contrast pair it belongs to, and a short rationale for the expected labels.
Provenance
Every specimen is original Fountain written for Inkwell and released under CC0-1.0.
How the labels work
The expected labels are line-level screenplay element types. A pair shares a pair_id, and split_group repeats it: keep both variants of a pair in the same split, because the point of a pair is the single difference between its two inputs.
Forty-two of the 48 specimens fall inside the subset a demonstration Fountain parser implements, and supported_by_demo marks which. The other six use dual dialogue, a note, boneyard, a section, lyrics and centered text — valid Fountain that the demonstration exporter stops on rather than guessing. The flag records one implementation's coverage.
Two more things the labels assume: an FDX paragraph can group adjacent action or dialogue lines, so a paragraph count can differ from the line-level label count; and a forced-element marker (!, @, ., >) is syntax, not rendered text.
What it's for
Teaching the syntax contrasts, building regression tests, and watching which single edit changes how a line is classified. Extend it with your own cases as your parser grows.
Load the records
from datasets import load_dataset
cases = load_dataset("Inkwell-Software/screenplay-format-edge-cases", split="test")
print(cases[0])Citation
Cite "Inkwell, Screenplay Format Edge Cases, version 1.0.0"; the machine-readable form is in `CITATION.cff`.
Maintenance
Open a discussion with the case ID, the expected behavior, the observed behavior, and the relevant part of the Fountain spec. A change to expected labels is a version increment and a changelog entry.
