mcmcmcmc/maia3-training-data
Maia3 Training Data Preprocessed chess position data for training Maia3 human move-prediction models. Positions are stored as precomputed features, not raw FENs, so training reads and expands them without re-tokenizing every epoch. Contents path rows size lichess_parquet/train_YYYY-MM_precomputed.parquet (31 files, 2023-01 … 2025-07) 334,438,119 ~33 GB allie_data/test_precomputed.parquet 884,049 69 MB allie_data/2022-test-annotated.jsonl — 45 MB… See the full description on the dataset page: https://huggingface.co/datasets/mcmcmcmc/maia3-training-data.
Maia3 Training Data
Preprocessed chess position data for training Maia3 human move-prediction models. Positions are stored as precomputed features, not raw FENs, so training reads and expands them without re-tokenizing every epoch.
Contents
Schema
All parquet files share one schema, with maia3_history: "7" in the file metadata (boards stored per position; a model may use fewer).
Boards are stored from the mover's perspective. move_idx is always present in legal_idx.
How it was produced
Training shards — monthly Lichess standard rated dumps, one parquet per month, via maia3-preprocess with --history 7 --balance. Game and position filtering happens at preprocess time; see maia3/preprocess.py in the repo. Reproduce with:
JOBS=3 data/download_and_rename_datasets.shValidation set — the ALLIE test set of Zhang et al. (2025): drop the opening plies, then truncate each game at the first position where the mover has less than the time threshold remaining. 2022-test-annotated.jsonl is the annotated source; the parquet is built from it with:
python data/allie_data/process_allie_data.py \
data/allie_data/2022-test-annotated.jsonl \
data/allie_data/test_precomputed.parquet --history 7Two validation sets
allie_data/test_precomputed.parquet is the primary one and what the configs point at. valid_2019-01_history_7_precomputed.parquet is a complement, not a replacement — same schema, so either loads with no code change:
Usage
hf download mcmcmcmc/maia3-training-data --repo-type dataset --local-dir dataThat lands the files where the training configs already expect them, so:
maia3-train --config configs/maia3-5m.yamlTo read a shard directly:
import pyarrow.parquet as pq
t = pq.read_table("lichess_parquet/train_2023-01_precomputed.parquet")License
Lichess game data is released under CC0 — these derived shards carry the same terms. The ALLIE test set follows the terms of Zhang et al. (2025).
