datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
old-games-transcript
90s Games Transcript
A small English-language corpus of narrative text from classic PC games.
This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading.
Dataset
Config
Records
Content
caesar3
20
Mission briefings + victory messages
diablo2_lod
7
Cinematic narration and dialogue
warcraft2
53
Human + Orc… See the full description on the dataset page: https://huggingface.co/datasets/sahar-millis-runi/old-games-transcript.transcript-formatter-curriculum
Transcript Formatter Curriculum (L0–L5 + RL)
Training data behind
Akash-Sakala/gpt-oss-120b-transcript-formatter-lora:
a layered curriculum that turns raw speech-to-text transcripts into clean,
formatted transcripts. Every row is input (raw transcript) → output
(formatted), across 21 categories spanning punctuation, casing, fillers,
disfluencies, homophones, ITN, proper nouns, URLs/emails, and layout.
Subsets (use the Data Viewer dropdown)
Each curriculum level is… See the full description on the dataset page: https://huggingface.co/datasets/Akash-Sakala/transcript-formatter-curriculum.huberman_transcript
