Team Ai
Datasetpublic

yrlyrl/spatial-mmcot-zebra_tetris

Spatial MMCoT v1 · zebra_tetris Zebra-CoT Tetris: polyomino puzzles of three kinds: apply a sequence of transformations to a shape, tile a shape with a set of pieces, and fill the grid outside a shape. The options are drawn in the input image, so the answer is a letter. Each step has its own target image: transformation puzzles start with a redraw of the start shape, then one image per transformation; tiling puzzles start with the isolated shape, fill-the-complement puzzles with… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-zebra_tetris.

sourceHugging Facecc-by-nc-4.0updated 9d agoView on Hugging Face
0likes452downloads
Dataset Card

Spatial MMCoT v1 · zebra_tetris

Zebra-CoT Tetris: polyomino puzzles of three kinds: apply a sequence of transformations to a shape, tile a shape with a set of pieces, and fill the grid outside a shape. The options are drawn in the input image, so the answer is a letter. Each step has its own target image: transformation puzzles start with a redraw of the start shape, then one image per transformation; tiling puzzles start with the isolated shape, fill-the-complement puzzles with the isolated shape and then its complement, and both then add one image per placed piece. All three kinds carry one task name, and no column tells them apart; the question's first words do (Apply the following sequence of transformations, Fill the exact <colour> shape, Fill the entire grid EXCEPT). The task table, the S13 balancing and the text-only baselines below pool the three kinds. The thoughts are upstream's text and were not checked against the images (see Known issues). Released rows carry 3 to 7 target images. The thoughts keep upstream's numbered labels ('THOUGHT 0:', 'THOUGHT 1:', ...) inside <think>. They count upstream's thoughts across the whole trace, not the row's slots, so one slot often holds two or more labelled thoughts, and a slot after a target image can open with the unlabelled end of the previous thought. No other source in this release has these labels; if you mix sources and do not want a model to learn them, remove them with re.sub(r'THOUGHT \d+:\s*', '', text). Refused by the conversion's checks: 1 row whose read-back is empty (S0.empty_readback); 2,440 rows whose target images are wrong -- a translation or scaling step pushed cells off the grid and upstream dropped them, so the labelled option is the clipped remnant, or a stated move drew no change (measured on the previous release by known_issues.py, replayed by rowuid from `blocklists/zebratetrisbadtargets.txt) (S5.knownbadtarget`).

Supervision kind (supervision_kind in meta): full_interleaved on every row: the upstream trace itself interleaves text and target images (drawn or rendered states on most sources; the source note above says which), and the read-back comes after the target image it reads. The text is upstream's and was not checked against the images.

Upstream: `multimodal-reasoning-lab/Zebra-CoT`. Licence: cc-by-nc-4.0. Changes from upstream: every image is decoded and re-encoded as JPEG, inputs at a 512 px long edge and targets at the size in meta.target_px (upstream's own size, at most 512 px; smaller on rows flagged S9.px=*); the image placeholders in the question and the reasoning were removed and the sentence around each repaired, sometimes by inserting words such as 'the image'; the reasoning is split into one thought per target; the ThinkMorph system prompt is prepended to the question; rows were filtered and balanced as described below.

Known issues

Rows with a measured per-row problem are listed in reports/known_issues/, one TSV per issue (a # <description> line, then row_uid<TAB>split<TAB>detail lines), so they can be filtered out. They are still in this release: no row was removed for these issues.

issuerowstrainvalidationwhathow it was foundfile
plan_names_answer_before_pieces4,9884,827161Tiling / fill-the-complement puzzle: the plan names the correct option before any of its pieces is drawn and no other option is tried; the piece images only confirm a choice already made in text.question starts 'Fill the exact' (tiling; first piece is target&#95;1) or 'Fill the entire grid EXCEPT' (complement; first piece is target&#95;2), and after removing enumerations the thoughts up to that piece's plan name 'option/set/choice &lt;answer&gt;'; measured 2026-10-01; known&#95;issues.py sha1 b2249331; export b3fa1b9 rows afa66fe1/92633acbreports/known_issues/plan_names_answer_before_pieces.tsv
shape_cell_count_contradicts_image1,0891,05237Tiling puzzle: the thought right after the isolated-shape image states a cell count for the shape that the image contradicts (almost always too few), so the row teaches a wrong read of an image the model has just drawn.question starts 'Fill the exact'; the first 'consists of / occupies / has ... N cells&#124;squares' in thought&#95;1 differs from the coloured cells counted in target&#95;0 (HSV S&gt;90, V&gt;70, cell-sized 4-connected components); measured 2026-10-01; known&#95;issues.py sha1 b2249331; export b3fa1b9 rows afa66fe1/92633acbreports/known_issues/shape_cell_count_contradicts_image.tsv
readback_names_no_option48246220The final thought never says which option matches (it only announces a comparison); the choice is only in &lt;answer&gt;.after removing enumerations ('options A, B, C and D', 'A-D'), the read-back names no option: no 'option/choice/set/piece/answer X', no 'is X.', no quoted '(X)'/'X', no 'the second option'; measured 2026-10-01; known&#95;issues.py sha1 b2249331; export b3fa1b9 rows afa66fe1/92633acbreports/known_issues/readback_names_no_option.tsv
text_says_options_not_shown27261The upstream text says the options are not shown or cites a trace, template or answer the model never sees ('options (not shown, but implied by the original trace)'), although every input image draws options A-D; training on it teaches the model to deny visible options.the question or any thought matches 'not shown&#124;not provided&#124;original trace&#124;in the trace&#124;implied by&#124;template indicates&#124;options are not&#124;inferred from the answer' (the review's pattern); measured 2026-10-01; known&#95;issues.py sha1 b2249331; export b3fa1b9 rows afa66fe1/92633acbreports/known_issues/text_says_options_not_shown.tsv
validation_target_identical_to_train909A target of this validation row is byte-identical to a training target (for some rows the puzzle's end state), so part of what the row asks the model to draw was seen in training.md5 of a validation target equals the md5 of some training target; measured 2026-10-01; known&#95;issues.py sha1 b2249331; export b3fa1b9 rows afa66fe1/92633acbreports/known_issues/validation_target_identical_to_train.tsv

To leave the listed rows out (the snippet in the loader section downloads reports/known_issues/ with the data):

python
import glob, os
root = "<root>/zebra_tetris"
drop = {line.split("\t")[0] for f in glob.glob(os.path.join(root, "reports/known_issues/*.tsv"))
        for line in open(f) if line.strip() and not line.startswith(("#", "row_uid\t"))}
# keep a row when its row_uid (a column of train/, meta/ and preview/) is not in drop

Measured caveats

Measured on this release by the pre-publication review (2026-09-25): problems that cannot be listed row by row (a shortcut in the options, a label convention, an upstream labelling scheme) and what the review found around the lists above. Where a caveat counts listed rows ("listed as ..."), the count is the table's, read from reports/known_issues/summary.json. Its other numbers are the review's own measurements, which no file carries: they hold for exactly these rows and are not re-measured automatically. Items marked Training-signal defect are problems in what the rows teach, not only in how they are described; no row was removed for them.

  • —Training-signal defect. In the tiling and fill-the-complement puzzles (4,994 rows, half the source) the plan names the correct option before any of its pieces is drawn, and only that option's pieces are placed: 4,988 rows (4,827 train, 161 validation; listed as plan_names_answer_before_pieces), 99.9% of those puzzles, name the labelled option at or before the plan of the first piece image (e.g. 6b42e76ba4799ec1: 'Option A contains three pieces. Let's test if these pieces can tile the red shape.'). The piece images confirm a choice the text has already made: use these rows as image supervision, not as evidence that generating the image helps. In the 2,559 transformation puzzles no option is named until the read-back after the last image.
  • —Training-signal defect. In tiling puzzles the thought after the isolated-shape image often states a cell count for the shape that the image contradicts, nearly always too low: at least 1,089 rows (1,052 train, 37 validation; listed as shape_cell_count_contradicts_image) (the list reads only 'consists of / occupies / has N cells' wordings), 43.7% of the 2,494 tiling rows and 14.4% of the source (e.g. 6b42e76ba4799ec1 says 'It occupies 14 cells' of a 25-cell shape). The piece sizes and totals quoted in the same thoughts are unreliable in the same way (2949258c0e73136b says '3, 3, and 3 cells, totaling 9'). These rows teach a wrong reading of the image the model has just drawn.
  • —Upstream's generator does not keep the shape whole at the grid edge: a translation or scaling drops the cells it would move off the grid, and a move that would push every cell off is not applied, so the image before and after it are identical. The previous release listed those rows (translation_clips_cells 2,242, scale_clips_cells 50, target_repeats_previous_target 541; 2,440 distinct rows); this release drops them at conversion (S5.known_bad_target, replayed from blocklists/zebra_tetris_bad_targets.txt), so none of the three lists has a row here. The transformation puzzles that remain are the 2,559 whose every step keeps the shape whole.
  • —482 rows (462 train, 20 validation; listed as readback_names_no_option), nearly all of them transformation puzzles (about one in five of that kind), end with a final thought that only says the final shape will be compared with the options and never says which option matches; the choice is then only in <answer>.
  • —The three puzzle kinds: transformation 2,559 rows (2,477 train, 82 validation), tiling 2,494 (2,411 / 83), fill-the-complement 2,500 (2,422 / 78).
  • —The 12 validation target slots with the content of a training target (see the split paragraph) lie in the 9 of the 243 validation rows listed as validation_target_identical_to_train; they are shapes other puzzles also reach (a redraw of the start shape, an intermediate step or the final result).

Size

splitrowstarget image slotsdistinct target images
train7,31035,19634,862
validation2431,1511,147

A slot is one target position in one row. Different puzzles pass through the same shapes, and a transformation that changes nothing visible (a mirror of a symmetric shape, a translation at the grid edge) repeats the previous image, so there are fewer distinct target images (by content hash) than slots.

tasktrainvalidation
tetris_transform7,310243

Input images per row: 1. Target images per row (the images the model is trained to generate): 3 to 7. Image corpus (source_scene_corpus): proceduralgrid 7,553. `proceduralgrid` is this pipeline's label for any procedurally generated synthetic image (2D grid puzzles and simple 3D renders of primitive objects alike); the source note above says what the images show.

Row format

One row is: input image(s) and a question, then K rounds of thought → target image (the target is the source's own ground-truth image, which the model is trained to generate), then a final thought (normally a read-back of the last target; where a source's final thought is something else, or often leaves out the answer, the source note or Known issues says so) and the answer; here K is 3 to 7. In the train config:

image_list        list<binary>  inputs first, then the K target images in order
num_input_images  int64         how many of image_list are inputs
instruction_list  list<string>  one element: system prompt + question; the options A-D are drawn in the input image, there is no option text
output_text_list  list<string>  K+1 elements:
  [0]   <think>plan 1</think><image_start>
  [j]   <image_end><think>plan j+1</think><image_start>
  [K]   <image_end><think>read-back</think><answer>answer</answer>
row_uid           string        join key to `meta` and `preview`

Every image is a JPEG, and no input image is larger than 512 px on its long edge (measured on this release, 2026-09-25); the size each target was stored at is target_px in meta.

<answer> holds exactly meta.answer_value (also the answer column of preview) on every row: score model output against that string.

The system prompt is ThinkMorph's VLM_THINK_SYSTEM_PROMPT from its inferencer.py, verbatim (GEN_THINK_SYSTEM_PROMPT there has the same text), including its leading and trailing newline. The markers are plain strings, not tokenizer special tokens; the prompt writes </image_end> and the data writes <image_end>, exactly as the ThinkMorph-7B checkpoint was trained.

preview shows the same rows with one column per slot: input_image_i for the inputs; for each of the K = num_steps rounds, the plan thought_j and its target target_image_j, both empty for K <= j < 7; and the read-back in thought_7 on every row.

meta holds the per-row sidecar: task, scene_id and geometry_uid (the scene and geometry keys; the split key is named in the split paragraph below), trajectory_id (a camera-path or sample label, empty where the source has none), num_steps, num_input_images, answer_type, answer_value, majority_class_rate, target_image_kind, target_px, est_tokens, licence, split (train / validation, the Hub split names), supervision_kind (full_interleaved / visual_aux / visual_only) and filter_flags. majority_class_rate is the share of the task's most frequent answer_value among its training rows: it measures answer skew and is not a guessing baseline (where a task mixes question types or each row has its own options it can be far below chance); compare scores with the text-only baselines below.

Per-row license in meta: cc-by-nc-4.0 7,553.

Flags on released rows (filter_flags in meta and preview, comma-separated):

flagrowsmeaning
S5.replay_unsupported7,553no solver re-derives this task's answer from the trace, so S5 did not replay it
S14.sampled_qa200chosen for the S14 human spot-check (reports/s14_sample.tsv)
S9.px=384112targets stored at this long-edge size (px) because the row was over the token budget at full size; the trainer's VAE transform (short edge at least 512) scales them back up, so these targets are trained as upsampled, softer images
S9.px=4489targets stored at this long-edge size (px) because the row was over the token budget at full size; the trainer's VAE transform (short edge at least 512) scales them back up, so these targets are trained as upsampled, softer images

Training with a BAGEL-family loader

Every row here has one input image (num_input_images is 1), so the stock ThinkMorph UnifiedEditIterableDataset (https://github.com/ThinkMorph/ThinkMorph: image_list[0] as input, image_list[j+1] after output_text_list[j]) and the IPT release's version (which reads num_input_images) both read it as intended. Mixed with a source whose rows have more than one input image, only a loader that reads num_input_images is correct.

The stock BAGEL edit loader (ByteDance-Seed/Bagel) cannot train these rows: it never reads output_text_list and expects each instruction_list element to be a list of paraphrases.

parquet_info.json keys each training chunk as <source>/<split>/<file>, here zebra_tetris/train/chunk_00000.parquet, with row-group counts read from the parquet footers. The loader matches a chunk only when its key equals the path it builds, os.path.join(data_dir, file), and skips a chunk with no key without a warning: a source that is alone in its group then fails with IndexError: list index out of range, and in a mixed group it adds no rows. Download into a directory named after the source, not after the repository:

python
from huggingface_hub import snapshot_download
snapshot_download("yrlyrl/spatial-mmcot-zebra_tetris", repo_type="dataset", local_dir="<root>/zebra_tetris",
                  allow_patterns=["train/*", "validation/*", "parquet_info.json", "reports/known_issues/*"])

Then either run from <root> with data_dir: zebra_tetris/train and parquet_info_path: zebra_tetris/parquet_info.json, or rebuild the index with absolute keys and use an absolute data_dir:

python
import json, os
root = "/abs/path/to/root"                      # the directory that holds zebra_tetris/
info = json.load(open(os.path.join(root, "zebra_tetris", "parquet_info.json")))
info = {os.path.join(root, k): v for k, v in info.items()}
json.dump(info, open(os.path.join(root, "zebra_tetris", "parquet_info_abs.json"), "w"))
# data_dir = os.path.join(root, "zebra_tetris", "train")   (spelled exactly so, no trailing slash)
# parquet_info_path = os.path.join(root, "zebra_tetris", "parquet_info_abs.json")

The Hugging Face cache (.../snapshots/<hash>/train/) or a folder named spatial-mmcot-zebra_tetris matches no key.

num_used_data counts chunk files, not rows: the loader repeats this source's file list up to that number, lists every (file, row group) pair, and deals whole row groups out, floor(R / worldsize) to each rank and floor(that / numworkers) to each DataLoader worker. The remainder is never read. This source has 1 training chunk file holding 58 row groups of up to 128 rows, so keep num_used_data large, e.g. the 128 of ThinkMorph's interleaved_reasoning.yaml (upstream's example.yaml asks for more than GPUs x workers); every row group is then read. Set to 1 and alone in its group on 8 GPUs with 4 workers, it reads only 32 of the 58 row groups. In a run that mixes sources, give each source the same multiple of its own training chunk-file count, e.g. 128 per file (128 here): the file list is repeated up to num_used_data entries, so a flat 128 for every source would read a two-file source's rows half as often as a one-file source's.

How the rows were chosen

stagerows
upstream rows read10,000
refused before conversion (S0raw)0
dropped at S5 (the text contains a phrase from S5's self-contradiction list, e.g. 'does not make sense', 'discrepancy', 'there must be a mistake'; a keyword match, not a comparison with the images, so it also removes some sound rows)2,445
refused by the final structural check (S0, after S9)1
after conversion and per-row filters7,554
removed by S10 (1 exact duplicate)1
removed by answer-prior balancing (S13)0
released7,553

Every removed row has one line, with its reason, in reports/:

filestepreason (the line's `flag`, or the field shown)rows
build/dropped.jsonlS0S0.empty_readback1
build/dropped.jsonlS5S5.known_bad_target2,440
build/dropped.jsonlS5S5.self_contradiction5
s10_dropped.jsonlS10level: exact1

<details><summary>Per-step counters of the conversion</summary>

10,000 upstream rows were read; S0raw refused 0 before a row existed and passed 10,000 to the first step. S0 runs once more, last, on the final bytes. The reason for every refused, dropped or quarantined row is in the files above.

stepinoutdroppedquarantinedrejectedrepaired
S410,00010,0000000
S4c10,00010,0000000
S510,0007,5552,445000
S87,5557,5550000
S97,5557,5550000
S0 (final structural check, after S9)7,5557,5541000

</details>

The train/validation split keeps rows sharing a scene_id in meta on one side, and the assignment is frozen (splits/ in the summary repository). scene_id is the md5 of the raw puzzle image, so a key is one puzzle. S12 saw 7,553 rows under 7,553 keys, one row per key, so the split is in effect per row. No validation input image has the content of a training input image, and none is a pixel-level near-copy of one. S12 does not record per source whether that test ran, but it skips it only for a source whose spec sets split_leak_pixels: false, and no spec does; over all sources it compared 21,661 candidate pairs (perceptual hash within 6 bits) pixel by pixel and found no near-copy (checked 2026-09-25). Target images are not covered by that check: 12 target-image slots in validation rows hold an image with the same content as a training target image (S12 compares a 64x64 greyscale hash and counts slots, so an image repeated in validation counts each time). These are states that other puzzles also reach, as a step or as their final state.

Answer-prior balancing (S13)

Each (task, split) group is checked separately. An answer is the answer value compared as lower-cased text without a trailing full stop, with 'farther' read as 'further' and 'nearer' as 'closer' (for multiple choice, the option text, not the letter; where the candidates are drawn in the image, as in zebrajigsaw and zebratetris, the answer is the letter itself). An answer is real when it holds at least 5 rows and 2% of the group; k is the number of real answers. Answer step: the target is max(30%, 1/k) when k >= 2, and max(30%, 1/d) over the d distinct answers when k = 1; a validation group uses the larger of its own target and its task's train target. A group is cut only when k >= 1 and its most common answer holds more than the target plus 5 percentage points; every answer is then capped at one common count, chosen so that none exceeds the target, and smaller answers keep all their rows. At the answer step, a group at or below that trigger, or with no real answer (k = 0), is left as it is, so its most common answer can hold up to the target plus 5 percentage points. A task whose train group has exactly two real answers is instead cut, in every split, so that its two largest answers have equal counts, with no trigger. Rank and label steps: then, in a group where every option value of every row is a number, the rank of the correct option among the sorted values, and after it, in a group where every trained answer is an option label, the label, are each capped by the same cut-and-trigger rule on their own counts (own target, validation included): capped, never evened out, so two labels are cut only when one exceeds 55%, and then only down to 50%. These steps can also cut groups the answer step left whole, including k = 0 groups, and can raise an answer's final share above its target; the run fails if a real answer ends above the target plus 5 percentage points. A train group of at least 20 rows in which one answer holds 90% or more fails the run. PET (exactcellspet) instead cuts each (question type x turn direction) cell to equal counts of its two answers; a PET cell that shows only one answer is removed.

tasksplitpassrulerows in → outreal answers ktargetlargest share, before → aftercut
tetris_transformtrainanswercap30[canon]7,310 → 7,310430.0%26.1% → 26.1%no
tetris_transformtrainlettercap30[letter]7,310 → 7,310430.0%26.1% → 26.1%no
tetris_transformvalidationanswercap30[canon]243 → 243430.0%27.6% → 27.6%no
tetris_transformvalidationlettercap30[letter]243 → 243430.0%27.6% → 27.6%no

S13 removed no row from this source.

Text-only baselines

Accuracy of guessers that never see an image. For each task the released training rows are split into two fixed halves by a hash of row_uid; each guesser is fitted on one half and scored once on the other (one held-out half, not cross-validation; eval rows below). The reference is chance (the mean of 1 / number of options) where every row is multiple choice, and otherwise the eval-half accuracy of always giving the answer most common in the fit half (when a task's top answers are nearly tied, this need not be the task's most common answer; the line after the table gives that answer's validation score). Accuracies are recounted from the stored rates and eval rows, so they are exact. A task is flagged when a text-only guesser beats its reference by more than 0.15 (for a free-form task, a guesser other than the most common answer). A flagged task can be partly answered from the text alone; an unflagged task passed only these probes, which do not prove the text carries no answer. Report scores on every task next to this baseline.

Guessers: keywords: the most common answer per set of spatial words in the question; last_mentioned: the option named last in the question body; letter_prior: the most common answer letter; majority: the answer most common in the fit half; option_prior: the option text that won most often when shown; template: the most common answer per question wording (numbers masked, object names kept).

taskbest text-only guesseraccuracyreferencemargineval rowsflagged
tetris_transformmajority0.2630.250 (chance)+0.0133,653no

Spot-check (S14)

Pending. The S14 rows are chosen and flagged S14.sampled_qa in meta and preview; the human pass over them has not been signed off yet.

Citation

Zebra-CoT is by Ang Li, Charles Wang, Deqing Fu, Kaiyu Yue, Zikui Cai, Wang Bill Zhu, Ollie Liu, Peng Guo, Willie Neiswanger, Furong Huang, Tom Goldstein and Micah Goldblum, released under CC BY-NC 4.0; this repository is a converted subset of it (the changes are listed under the licence line above). Its card asks users to cite:

bibtex
@inproceedings{li2026zebracot,
  title={Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning},
  author={Ang Li and Charles Wang and Deqing Fu and Kaiyu Yue and Zikui Cai and Wang Bill Zhu and Ollie Liu and Peng Guo and Willie Neiswanger and Furong Huang and Tom Goldstein and Micah Goldblum},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
  url={https://openreview.net/forum?id=c6XIVI3TiQ}
}

Provenance

The release files were written by our conversion code (the code repository is not public yet), scripts/convert/export.py at commit b3fa1b993ca9, from build zebra_tetris_r3. The build was made by scripts/convert/run_source.py from the same repository at commit 76c78bd817c2. S10, S12 and S13 ran before the export; reports/export_manifest.json pins every input the export read by SHA-1 (build_manifest_sha1, s10_keep_sha1, s12_assignments_sha1, s13_balanced_keep_sha1).

Every row removed between upstream and this release has one line, with its reason, in reports/: build/dropped.jsonl (rows refused before conversion or dropped by a conversion step); build/quarantine.jsonl (rows set aside by S4c because an automatic check could not match the read-back's conclusion to the label); s10_dropped.jsonl (duplicates removed by S10); s10_label_conflicts.jsonl (rows S10 withheld because another row asks the identical question, options in the same order, of the same images with a different answer); s13_dropped.jsonl (rows removed by answer-prior balancing). known_issues/ lists rows with a measured problem (see Known issues); reports/ also holds the build manifest (absolute paths cut to basenames) and counters, the S14 sample list (s14_sample.tsv: rowuid, task, split) and `exportmanifest.json. Part of [yrlyrl/spatial-mmcot`](https://huggingface.co/datasets/yrlyrl/spatial-mmcot).