Team Ai
Datasetpublic

KRAFTON/A2Z-GameSpec-Bench

A2Z GameSpec-Bench A2Z GameSpec-Bench contains 100 long-form Game Design Documents (GDDs) for evaluating how faithfully coding agents implement game specifications. Each task pairs a GDD with a Dependency-Aware Contract and canonical scenarios. The benchmark combines source-code inspection, scenario-based replay, and adaptive playtesting with agent-generated Test Policies. Paper · Project page · Code repository Dataset There are 50 Small tasks and 50 Big… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/A2Z-GameSpec-Bench.

sourceHugging Facemitupdated 2d agoView on Hugging Face
2likes420downloads
Dataset Card

<div align="center"> <img src="assets/header.png" width="100%" alt="A2Z GameSpec-Bench"> </div>

A2Z GameSpec-Bench

A2Z GameSpec-Bench contains 100 long-form Game Design Documents (GDDs) for evaluating how faithfully coding agents implement game specifications. Each task pairs a GDD with a Dependency-Aware Contract and canonical scenarios. The benchmark combines source-code inspection, scenario-based replay, and adaptive playtesting with agent-generated Test Policies.

Paper · Project page · Code repository

<div align="center"> <img src="assets/teaser_intro.gif" width="100%" alt="A2Z GameSpec-Bench introduction"> </div>

Dataset

There are 50 Small tasks and 50 Big tasks. Small tasks focus on compact game designs; Big tasks specify more extensive, interacting systems. Both are evaluation tasks. The default configuration contains all 100 examples in the test split; the small and big configurations each contain 50 examples. Task size reflects the scope of the game design.

This repository provides the benchmark GDDs, contracts, and canonical scenarios.

SmallBig
Games5050
Rules8691,860
Invariants1,3781,862
Canonical scenarios6461,385
text
data/                         Small and Big Parquet tables
gdds/{small,big}/              Original GDD Markdown files
contracts/{small,big}/<game>/  One contract and scenario set per task
library/{small,big}/<game>/    PNG art assets for the games
metadata/                     File hashes and counts calculated from this release

Each contract directory contains:

  • —rules.json: causal rules with preconditions, a trigger, effects, and GDD provenance. Dependency relations can be derived from matching state/event fields.
  • —invariants.json: scoped constraints and explicit initialization requirements. Initialization requirements include their applicable event and scope.
  • —scenarios.json: canonical runtime situations, setup descriptions, and observations to check. exposes lists the behaviors covered by each scenario.
  • —rules.vocab.json: entity, attribute, and event names used by the contract, including trigger events.
  • —requirements.json: source requirement entries and their corresponding graph-rule IDs.
  • —state_sources.json: source passages describing state variables and explicitly identified initialization cases.

Source references contain a repository-relative file, 1-based inclusive line_start/line_end, and the exact original text in quote. Abbreviated or combined scenario descriptions, where retained, are called source_summary; the accompanying source references contain the verbatim GDD passages. Structured values may contain an opaque description or expression when the requirement is naturally expressed in text. These descriptions preserve requirements that need semantic interpretation.

Usage

python
import json
from datasets import load_dataset

# All 100 tasks
dataset = load_dataset("KRAFTON/A2Z-GameSpec-Bench", split="test")

# Small or Big tasks only
small = load_dataset("KRAFTON/A2Z-GameSpec-Bench", "small", split="test")
big = load_dataset("KRAFTON/A2Z-GameSpec-Bench", "big", split="test")

game = dataset[0]
gdd = game["gdd"]
contract = json.loads(game["contract_json"])
rules = contract["rules"]
invariants = contract["invariants"]

The Parquet table stores one game per row. contract_json preserves heterogeneous structured values without Arrow type coercion; parsing it yields the same data as the six JSON files in the corresponding contract directory. No custom loading script or remote code is required.

Other columns are game_id, task_size, title, gdd, gdd_tokens, rule_count, invariant_count, scenario_count, and source_requirement_count. Counts are computed directly from the included files.

Evaluation environment

The main benchmark targets 2D single-player browser games using a common Phaser 4.1, TypeScript, and Vite environment. Preserve the intended game behavior when adapting platform-specific implementation details. The evaluation contract is withheld during initial game generation; canonical scenarios support fixed replay preparation. Keep the selected release's contracts fixed across models and revision rounds.

Citation

bibtex
@article{lee2026a2zgamespecbench,
  title={A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?},
  author={Lee, Seonho and Jeong, Wonryeol and Cereser, Alberto and Kang, Inha and Kim, Hyeonjong and Kwak, Seungmin and Park, Dongmin},
  journal={arXiv preprint arXiv:2609.39564},
  year={2026},
  url={https://arxiv.org/abs/2609.39564}
}

License

Released under the MIT License. Copyright © 2026 KRAFTON, Inc.