Team Ai
Datasetpublic

Inkwell-Software/screenplay-revision-evaluation

Screenplay Revision Evaluation Cases 24 original screenwriting revision tasks. Each gives a short scene and a constraint — cut a page to its beat, plant a prop, hold an answer back, fix a continuity slip — then pairs it with mechanical checks (a word ceiling, a line that must survive) and separate human-review questions. It tests whether a tool, or a person, can make a tightly-constrained edit while keeping the scene intact. Each task's reference_output is null, because a… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-revision-evaluation.

sourceHugging Facecc0-1.0updated 15d agoView on Hugging Face
0likes53downloads
Dataset Card

Screenplay Revision Evaluation Cases

24 original screenwriting revision tasks. Each gives a short scene and a constraint — cut a page to its beat, plant a prop, hold an answer back, fix a continuity slip — then pairs it with mechanical checks (a word ceiling, a line that must survive) and separate human-review questions. It tests whether a tool, or a person, can make a tightly-constrained edit while keeping the scene intact. Each task's reference_output is null, because a creative edit has more than one right version.

Version: 1.0.0 · Maintainer: Inkwell.

Built by Inkwell, the IDE for screenwriters. For a related continuity exercise, reorder seven scene cards and inspect their setups.

What's included

data/cases.jsonl holds 24 records; schema.json is a JSON Schema for one record. Each record carries the scene input, the revision constraint, the mechanical checks, and the human-review questions.

Provenance

Every case is original, written for Inkwell, and released under CC0-1.0.

Task design

The 24 cases cover information boundaries, prop and spatial continuity, refusals, visible deadlines, delayed answers, scene exits, and other explicit revision constraints. Each has an original short input and an instruction.

The test split is a publication convention: once you tune prompts on these cases they become development examples, so keep a held-out set for comparison claims.

Reading the checks

The word ceiling, exact anchor, and introductory-preface heuristic are mechanical checks — they measure whether the constraint was met, and the human-review questions cover dramatic quality and meaning. A reviewer explains failures and trade-offs rather than collapsing unlike judgments into one score.

What it's for

Regression-test an editing workflow, develop a task-specific rubric, or teach constrained revision. It is a focused authored set of 24 tasks — extend it with your own as your rubric grows.

Load the records

python
from datasets import load_dataset
cases = load_dataset("Inkwell-Software/screenplay-revision-evaluation", split="test")
print(cases[0])

Citation

Cite "Inkwell, Screenplay Revision Evaluation Cases, version 1.0.0"; the machine-readable form is in `CITATION.cff`.

Maintenance

Open a discussion with the case ID, the constraint, and what a tool or reviewer produced. A change to a case's input, constraint, or checks is a version increment and a changelog entry.