Team Ai
Datasetpublic

sahar-millis-runi/old-games-transcript

90s Games Transcript A small English-language corpus of narrative text from classic PC games. This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading. Dataset Config Records Content caesar3 20 Mission briefings + victory messages diablo2_lod 7 Cinematic narration and dialogue warcraft2 53 Human + Orc… See the full description on the dataset page: https://huggingface.co/datasets/sahar-millis-runi/old-games-transcript.

sourceHugging Facemitupdated 13d agoView on Hugging Face
1likes84downloads
Dataset Card

90s Games Transcript

A small English-language corpus of narrative text from classic PC games.

This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading.

<img src="https://cdn-uploads.huggingface.co/production/uploads/6a9ec4fa322229fd8f3c2b3c/ZFOtUY8kEL2YZcUX7Wi_z.png" alt="logo" width="700">

Dataset

ConfigRecordsContent
caesar320Mission briefings + victory messages
diablo2_lod7Cinematic narration and dialogue
warcraft253Human + Orc mission briefings and ending text
Total80~11.8k words

Each transcript has its own schema.

Quick start

python
from datasets import load_dataset

caesar3 = load_dataset(
    "user_namei/old-games-transcript",
    "caesar3",
    split="train",
)

diablo2 = load_dataset(
    "user_namei/old-games-transcript",
    "diablo2_lod",
    split="train",
)

warcraft2 = load_dataset(
    "user_namei/old-games-transcript",
    "warcraft2",
    split="train",
)

The train split is simply a Hugging Face loading convention. This dataset is not organized as an ML train/test benchmark.

What's inside?

Caesar III

The caesar3 configuration contains campaign assignments from both the Peaceful and Military paths.

Each record includes metadata such as mission, rank, assignment, and map, together with:

  • —brief_transcript — the pre-mission briefing
  • —win_transcript — the post-victory message

This makes the game particularly interesting for studying the relationship between instruction, player performance, evaluation, and progression.

Diablo II

The diablo2_lod configuration contains seven cinematic records, from The Sister's Lament through Destruction's End.

Each row represents a complete cinematic rather than individual speaker turns. The transcripts therefore preserve the scene as the primary unit of analysis.

act is a dataset sequence identifier and should not necessarily be interpreted as canonical in-game act numbering.

Warcraft II

The warcraft2 configuration contains Human and Orc material from:

  • —Warcraft II: Tides of Darkness (TOD)
  • —Warcraft II: Beyond the Dark Portal (BTDP)

The corpus primarily consists of mission briefings and is well suited to examining military instruction, factional framing, player address, objectives, and narrative justification of gameplay.

Research uses

Potential applications include game studies, discourse analysis, player agency and address, narrative temporality, factional rhetoric, computational stylistics, and historical analysis of game writing.

This is a small, curated corpus, so quantitative findings are best used to guide and support close reading rather than to make broad claims about video-game language in general.

Audio sources

Corresponding game audio is available through the Internet Archive: LINK1, LINK2, LINK3

The audio is not included in this dataset. These links are provided as reference material and can be used to verify transcripts or study features that text alone cannot represent, such as performance, timing, pronunciation, and prosody.

Acknowledgements

This dataset exists because of the original writers, designers, performers, and developers behind these games, as well as the preservation efforts that keep historical game media accessible.

Special thanks to the fan communities that continue to document, preserve, and discuss these games decades after their release:

And a small shoutout to the fans of Warcraft II, Caesar III, and Diablo II who are still playing, modding, documenting, discussing, and preserving these games. Long-lived fan communities are a major reason pieces of game history like these remain accessible.

All game titles and trademarks belong to their respective owners.

Copyright and usage

The transcripts reproduce dialogue and text from copyrighted commercial video games. The underlying game text, characters, performances, trademarks, and other source material remain the property of their respective rights holders.

This dataset does not grant or imply a license to the underlying game content. It is provided as a research and documentation resource, and users are responsible for determining whether their intended use is permitted under applicable law.

No game executables, installation media, or audio files are included.