Team Ai
Datasetpublic

mcmcmcmc/maia3-training-data

Maia3 Training Data Preprocessed chess position data for training Maia3 human move-prediction models. Positions are stored as precomputed features, not raw FENs, so training reads and expands them without re-tokenizing every epoch. Contents path rows size lichess_parquet/train_YYYY-MM_precomputed.parquet (31 files, 2023-01 … 2025-07) 334,438,119 ~33 GB allie_data/test_precomputed.parquet 884,049 69 MB allie_data/2022-test-annotated.jsonl — 45 MB… See the full description on the dataset page: https://huggingface.co/datasets/mcmcmcmc/maia3-training-data.

sourceHugging Facecc0-1.0updated 16d agoView on Hugging Face
1likes321downloads
Dataset Card

Maia3 Training Data

Preprocessed chess position data for training Maia3 human move-prediction models. Positions are stored as precomputed features, not raw FENs, so training reads and expands them without re-tokenizing every epoch.

Contents

pathrowssize
lichess_parquet/train_YYYY-MM_precomputed.parquet (31 files, 2023-01 … 2025-07)334,438,119~33 GB
allie_data/test_precomputed.parquet884,04969 MB
allie_data/2022-test-annotated.jsonl—45 MB
valid_2019-01_history_7_precomputed.parquet1,000,000104 MB

Schema

All parquet files share one schema, with maia3_history: "7" in the file metadata (boards stored per position; a model may use fewer).

columntype
piece_idslist<list<uint8>>per-board piece ids, 64 squares per board, oldest board first
castlinglist<uint8>castling rights per board, bit-packed
ep_squarelist<int8>en-passant target square per board, -1 if none
move_idxint16index of the played move in the fixed 4352-move vocabulary
legal_idxlist<int16>indices of all legal moves in that vocabulary
self_eloint32Elo of the side to move (900–2600)
oppo_eloint32Elo of the opponent (900–2600)
value_targetint8game result from the mover's view — 0 loss, 1 draw, 2 win

Boards are stored from the mover's perspective. move_idx is always present in legal_idx.

How it was produced

Training shards — monthly Lichess standard rated dumps, one parquet per month, via maia3-preprocess with --history 7 --balance. Game and position filtering happens at preprocess time; see maia3/preprocess.py in the repo. Reproduce with:

bash
JOBS=3 data/download_and_rename_datasets.sh

Validation set — the ALLIE test set of Zhang et al. (2025): drop the opening plies, then truncate each game at the first position where the mover has less than the time threshold remaining. 2022-test-annotated.jsonl is the annotated source; the parquet is built from it with:

bash
python data/allie_data/process_allie_data.py \
  data/allie_data/2022-test-annotated.jsonl \
  data/allie_data/test_precomputed.parquet --history 7

Two validation sets

allie_data/test_precomputed.parquet is the primary one and what the configs point at. valid_2019-01_history_7_precomputed.parquet is a complement, not a replacement — same schema, so either loads with no code change:

ALLIE test set2019-01 held-out month
positions884,0491,000,000
source month(s)20222019-01
relative to training data (2023-01 … 2025-07)beforewell before
clock filteringopenings dropped, truncated at low timenone
measuresin-distribution move predictiongeneralization across time

Usage

bash
hf download mcmcmcmc/maia3-training-data --repo-type dataset --local-dir data

That lands the files where the training configs already expect them, so:

bash
maia3-train --config configs/maia3-5m.yaml

To read a shard directly:

python
import pyarrow.parquet as pq
t = pq.read_table("lichess_parquet/train_2023-01_precomputed.parquet")

License

Lichess game data is released under CC0 — these derived shards carry the same terms. The ALLIE test set follows the terms of Zhang et al. (2025).