sahar-millis-runi/old-games-transcript
90s Games Transcript A small English-language corpus of narrative text from classic PC games. This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading. Dataset Config Records Content caesar3 20 Mission briefings + victory messages diablo2_lod 7 Cinematic narration and dialogue warcraft2 53 Human + Orc… See the full description on the dataset page: https://huggingface.co/datasets/sahar-millis-runi/old-games-transcript.
90s Games Transcript
A small English-language corpus of narrative text from classic PC games.
This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading.
<img src="https://cdn-uploads.huggingface.co/production/uploads/6a9ec4fa322229fd8f3c2b3c/ZFOtUY8kEL2YZcUX7Wi_z.png" alt="logo" width="700">
Dataset
Each transcript has its own schema.
Quick start
from datasets import load_dataset
caesar3 = load_dataset(
"user_namei/old-games-transcript",
"caesar3",
split="train",
)
diablo2 = load_dataset(
"user_namei/old-games-transcript",
"diablo2_lod",
split="train",
)
warcraft2 = load_dataset(
"user_namei/old-games-transcript",
"warcraft2",
split="train",
)The train split is simply a Hugging Face loading convention. This dataset is not organized as an ML train/test benchmark.
What's inside?
Caesar III
The caesar3 configuration contains campaign assignments from both the Peaceful and Military paths.
Each record includes metadata such as mission, rank, assignment, and map, together with:
brief_transcript— the pre-mission briefingwin_transcript— the post-victory message
This makes the game particularly interesting for studying the relationship between instruction, player performance, evaluation, and progression.
Diablo II
The diablo2_lod configuration contains seven cinematic records, from The Sister's Lament through Destruction's End.
Each row represents a complete cinematic rather than individual speaker turns. The transcripts therefore preserve the scene as the primary unit of analysis.
act is a dataset sequence identifier and should not necessarily be interpreted as canonical in-game act numbering.
Warcraft II
The warcraft2 configuration contains Human and Orc material from:
- Warcraft II: Tides of Darkness (
TOD) - Warcraft II: Beyond the Dark Portal (
BTDP)
The corpus primarily consists of mission briefings and is well suited to examining military instruction, factional framing, player address, objectives, and narrative justification of gameplay.
Research uses
Potential applications include game studies, discourse analysis, player agency and address, narrative temporality, factional rhetoric, computational stylistics, and historical analysis of game writing.
This is a small, curated corpus, so quantitative findings are best used to guide and support close reading rather than to make broad claims about video-game language in general.
Audio sources
Corresponding game audio is available through the Internet Archive: LINK1, LINK2, LINK3
The audio is not included in this dataset. These links are provided as reference material and can be used to verify transcripts or study features that text alone cannot represent, such as performance, timing, pronunciation, and prosody.
Acknowledgements
This dataset exists because of the original writers, designers, performers, and developers behind these games, as well as the preservation efforts that keep historical game media accessible.
Special thanks to the fan communities that continue to document, preserve, and discuss these games decades after their release:
- Warcraft II: Warcraft Wiki — Tides of Darkness · Warcraft Wiki — Beyond the Dark Portal
- Caesar III: City-Building Games Wiki — Caesar III
- Diablo II: Diablo Wiki — Lord of Destruction
And a small shoutout to the fans of Warcraft II, Caesar III, and Diablo II who are still playing, modding, documenting, discussing, and preserving these games. Long-lived fan communities are a major reason pieces of game history like these remain accessible.
All game titles and trademarks belong to their respective owners.
Copyright and usage
The transcripts reproduce dialogue and text from copyrighted commercial video games. The underlying game text, characters, performances, trademarks, and other source material remain the property of their respective rights holders.
This dataset does not grant or imply a license to the underlying game content. It is provided as a research and documentation resource, and users are responsible for determining whether their intended use is permitted under applicable law.
No game executables, installation media, or audio files are included.
