Inkwell-Software/dialogue-word-concentration
Dialogue Word Concentration A reproducible numerical analysis of the Cornell Movie-Dialogs Corpus: 617 movie IDs, 304,439 word-bearing sampled utterances, 3,210,011 words. It carries per-film measurements and source-matched titles, credited to Cornell. Built by Inkwell, the IDE for screenwriters. Change the threshold and inspect the distribution across films in the interactive explorer, or read the dialogue methods and sources. What the files hold data/films.csv:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/dialogue-word-concentration.
Dialogue Word Concentration
A reproducible numerical analysis of the Cornell Movie-Dialogs Corpus: 617 movie IDs, 304,439 word-bearing sampled utterances, 3,210,011 words. It carries per-film measurements and source-matched titles, credited to Cornell.
Built by Inkwell, the IDE for screenwriters. Change the threshold and inspect the distribution across films in the interactive explorer, or read the dialogue methods and sources.
What the files hold
data/films.csv: 617 rows, one per movie ID — source-matched title, sampled utterance and word totals, Gini, and the word share carried by each film's longest ceil(10% × n) utterances.
data/thresholds.csv: six rows, one per word-length threshold, with pooled and equal-film proportions and 95% film-bootstrap intervals. It loads as its own config, one row per threshold.
film-histograms.json: per-film length histograms and sampled Lorenz points. summary.json: denominators, exclusion policy, source hash, and bootstrap settings.
Findings
The longest 30,444 word-bearing utterances (ceil(10% × 304,439)) hold 35.7% of the words. Utterances over 30 words are 5.4% of the sample but 24.3% of the words — 22.4% when each film counts equally. Of the 304,713 projected rows, 274 have zero words; those sit in the input audit but stay out of the lexical-length denominators. Exact values are in summary.json and thresholds.csv.
These are descriptive statistics over sampled dialogue — how words distribute across each film's turns.
Methods
For each film: sort positive word counts, compute total words, the Lorenz curve and Gini, and the share carried by the longest ceil(10% × n) observations (equal-length ties do not expand the selected count). Pooled shares divide pooled numerators by pooled denominators; equal-film shares average each film's own proportion. The bootstrap resamples 617 films with replacement 5,000 times and reports the 2.5th–97.5th percentile of the equal-film mean, describing resampling stability within this convenience sample (seed 20260917, recorded in summary.json).
Source and attribution
Cristian Danescu-Niculescu-Mizil and Lillian Lee, "Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs," 2011 (arXiv:1106.3077). Corpus page. Cite this dataset and that paper; both are in `CITATION.cff`.
Load the data
from datasets import load_dataset
films = load_dataset("Inkwell-Software/dialogue-word-concentration", "films", split="sample")
thresholds = load_dataset("Inkwell-Software/dialogue-word-concentration", "thresholds", split="sample")
assert len(films) == 617Version and maintenance
Version 1.0.0. Reports should state the version and the input hash in summary.json; a correction preserves the prior methodology record and explains any change to denominators.
