KRAFTON/A2Z-GameSpec-Bench
A2Z GameSpec-Bench A2Z GameSpec-Bench contains 100 long-form Game Design Documents (GDDs) for evaluating how faithfully coding agents implement game specifications. Each task pairs a GDD with a Dependency-Aware Contract and canonical scenarios. The benchmark combines source-code inspection, scenario-based replay, and adaptive playtesting with agent-generated Test Policies. Paper · Project page · Code repository Dataset There are 50 Small tasks and 50 Big… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/A2Z-GameSpec-Bench.
<div align="center"> <img src="assets/header.png" width="100%" alt="A2Z GameSpec-Bench"> </div>
A2Z GameSpec-Bench
A2Z GameSpec-Bench contains 100 long-form Game Design Documents (GDDs) for evaluating how faithfully coding agents implement game specifications. Each task pairs a GDD with a Dependency-Aware Contract and canonical scenarios. The benchmark combines source-code inspection, scenario-based replay, and adaptive playtesting with agent-generated Test Policies.
Paper · Project page · Code repository
<div align="center"> <img src="assets/teaser_intro.gif" width="100%" alt="A2Z GameSpec-Bench introduction"> </div>
Dataset
There are 50 Small tasks and 50 Big tasks. Small tasks focus on compact game designs; Big tasks specify more extensive, interacting systems. Both are evaluation tasks. The default configuration contains all 100 examples in the test split; the small and big configurations each contain 50 examples. Task size reflects the scope of the game design.
This repository provides the benchmark GDDs, contracts, and canonical scenarios.
data/ Small and Big Parquet tables
gdds/{small,big}/ Original GDD Markdown files
contracts/{small,big}/<game>/ One contract and scenario set per task
library/{small,big}/<game>/ PNG art assets for the games
metadata/ File hashes and counts calculated from this releaseEach contract directory contains:
rules.json: causal rules with preconditions, a trigger, effects, and GDD provenance. Dependency relations can be derived from matching state/event fields.invariants.json: scoped constraints and explicit initialization requirements. Initialization requirements include their applicable event and scope.scenarios.json: canonical runtime situations, setup descriptions, and observations to check.exposeslists the behaviors covered by each scenario.rules.vocab.json: entity, attribute, and event names used by the contract, including trigger events.requirements.json: source requirement entries and their corresponding graph-rule IDs.state_sources.json: source passages describing state variables and explicitly identified initialization cases.
Source references contain a repository-relative file, 1-based inclusive line_start/line_end, and the exact original text in quote. Abbreviated or combined scenario descriptions, where retained, are called source_summary; the accompanying source references contain the verbatim GDD passages. Structured values may contain an opaque description or expression when the requirement is naturally expressed in text. These descriptions preserve requirements that need semantic interpretation.
Usage
import json
from datasets import load_dataset
# All 100 tasks
dataset = load_dataset("KRAFTON/A2Z-GameSpec-Bench", split="test")
# Small or Big tasks only
small = load_dataset("KRAFTON/A2Z-GameSpec-Bench", "small", split="test")
big = load_dataset("KRAFTON/A2Z-GameSpec-Bench", "big", split="test")
game = dataset[0]
gdd = game["gdd"]
contract = json.loads(game["contract_json"])
rules = contract["rules"]
invariants = contract["invariants"]The Parquet table stores one game per row. contract_json preserves heterogeneous structured values without Arrow type coercion; parsing it yields the same data as the six JSON files in the corresponding contract directory. No custom loading script or remote code is required.
Other columns are game_id, task_size, title, gdd, gdd_tokens, rule_count, invariant_count, scenario_count, and source_requirement_count. Counts are computed directly from the included files.
Evaluation environment
The main benchmark targets 2D single-player browser games using a common Phaser 4.1, TypeScript, and Vite environment. Preserve the intended game behavior when adapting platform-specific implementation details. The evaluation contract is withheld during initial game generation; canonical scenarios support fixed replay preparation. Keep the selected release's contracts fixed across models and revision rounds.
Citation
@article{lee2026a2zgamespecbench,
title={A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?},
author={Lee, Seonho and Jeong, Wonryeol and Cereser, Alberto and Kang, Inha and Kim, Hyeonjong and Kwak, Seungmin and Park, Dongmin},
journal={arXiv preprint arXiv:2609.39564},
year={2026},
url={https://arxiv.org/abs/2609.39564}
}License
Released under the MIT License. Copyright © 2026 KRAFTON, Inc.
