yrlyrl/spatial-mmcot-zebra_jigsaw
Spatial MMCoT v1 · zebra_jigsaw Zebra-CoT visual jigsaw, type-2 (single-image) rows only, on ImageNet images (mostly photographs; some are web graphics such as banners and flyer templates). Each row has one input image: the source image with its missing piece(s) greyed out, above a panel of four candidate piece sets labelled A-D. The options exist only in that image, so the answer is the letter. The target is the complete source image. Type-1 rows are excluded: three quarters of… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-zebra_jigsaw.
Spatial MMCoT v1 · zebra_jigsaw
Zebra-CoT visual jigsaw, type-2 (single-image) rows only, on ImageNet images (mostly photographs; some are web graphics such as banners and flyer templates). Each row has one input image: the source image with its missing piece(s) greyed out, above a panel of four candidate piece sets labelled A-D. The options exist only in that image, so the answer is the letter. The target is the complete source image. Type-1 rows are excluded: three quarters of their target images are deliberately wrong assemblies, and their text leaks the answer. The input is upstream's 2400x1800 puzzle sheet downscaled to 512x384, so each candidate piece is about 4.7 times smaller per side than upstream; the targets keep their ImageNet size (mostly 500 px on the long edge). On some rows the upstream final thought only restates the task and never says which option it picked; the choice is then only in <answer>. The thoughts keep upstream's numbered labels ('THOUGHT 0:', 'THOUGHT 1:', ...) inside <think>. They count upstream's thoughts across the whole trace, not the row's slots, so one slot often holds two or more labelled thoughts, and a slot after a target image can open with the unlabelled end of the previous thought. No other source in this release has these labels; if you mix sources and do not want a model to learn them, remove them with re.sub(r'THOUGHT \d+:\s*', '', text).
Reasoning text in this release (r4)
The paragraph above describes the upstream data, and where it describes the reasoning text it describes the text of the previous release (r3). In this release every thought was rewritten: Qwen3.8-27B in thinking mode (temperature 1.0, top-p 0.95, top-k 20; prompt v4.1, scripts/annotate/rationale.py in our conversion code (the code repository is not public yet)) was shown the row's input images, its target images in order, the question, the options and the answer, and wrote the thought before each target image (what is still unknown, and what the picture should show to settle it) and the read-back after the last one. The images, the question, the options, the answer and the number of steps are r3's, checked row by row (image bytes, instruction, K and the <answer> value).
This is the training subset: 108 rows of r4 are not in it, listed with their reason in reports/subset_dropped.jsonl: 108 the rewrite model judged the stored answer doubtful (its note is in the list) (label_doubt).
Self-check of the rewritten text (measured 2026-10-02 on 22 rows sampled from this source before the removals above, each judged against its full-resolution pictures): 0.86 faithfulness errors per row (a claim the pictures contradict), 32% of rows with none; the plan names the real obstacle on 100%, the read-back reads the picture it follows on 100%, and the answer follows from the text on 100%.
The measured caveats on the r3 card were counted on r3's rows and text and are not repeated here; the per-row known-issue lists below were measured again on these rows and this text.
Supervision kind (supervision_kind in meta): full_interleaved on every row: the upstream trace itself interleaves text and target images (drawn or rendered states on most sources; the source note above says which), and the read-back comes after the target image it reads. In this release the text was rewritten by a model shown every image of the row (see the r4 section above).
Upstream: `multimodal-reasoning-lab/Zebra-CoT`. Licence: cc-by-nc-4.0. This is the licence Zebra-CoT declares. The images are ImageNet images: copyright stays with their original rights holders (mostly photographers, some of whom are credited in the image itself; for the web graphics, the companies or designers who made them), and ImageNet's terms of access (https://image-net.org/download.php: non-commercial research and educational use only, and passing the images on only to people who agree to those terms) apply as well. Changes from upstream: every image is decoded and re-encoded as JPEG, the input at 512x384 and the target at its ImageNet size (at most 1,024 px); the image placeholders in the question and the reasoning were removed and the sentence around each repaired, sometimes by inserting words such as 'the image'; the reasoning is split into one thought per target; the ThinkMorph system prompt is prepended to the question; rows were filtered and balanced as described below.
Known issues
Rows with a measured per-row problem are listed in reports/known_issues/, one TSV per issue (a # <description> line, then row_uid<TAB>split<TAB>detail lines), so they can be filtered out. They are still in this release: no row was removed for these issues.
To leave the listed rows out (the snippet in the loader section downloads reports/known_issues/ with the data):
import glob, os
root = "<root>/zebra_jigsaw"
drop = {line.split("\t")[0] for f in glob.glob(os.path.join(root, "reports/known_issues/*.tsv"))
for line in open(f) if line.strip() and not line.startswith(("#", "row_uid\t"))}
# keep a row when its row_uid (a column of train/, meta/ and preview/) is not in dropSize
A slot is one target position in one row. Some training rows are different puzzles cut from the same image: their inputs differ but their target, the complete image, is the same (they share a geometry_uid, not a scene_id), so there are fewer distinct target images (by content hash) than slots.
Input images per row: 1. Target images per row (the images the model is trained to generate): 1. Image corpus (source_scene_corpus): imagenet 10,854.
Row format
One row is: input image(s) and a question, then K rounds of thought → target image (the target is the source's own ground-truth image, which the model is trained to generate), then a final thought (normally a read-back of the last target; where a source's final thought is something else, or often leaves out the answer, the source note or Known issues says so) and the answer; here K is 1. In the train config:
image_list list<binary> inputs first, then the K target images in order
num_input_images int64 how many of image_list are inputs
instruction_list list<string> one element: system prompt + question; the options A-D are drawn in the input image, there is no option text
output_text_list list<string> K+1 elements:
[0] <think>plan 1</think><image_start>
[j] <image_end><think>plan j+1</think><image_start>
[K] <image_end><think>read-back</think><answer>answer</answer>
row_uid string join key to `meta` and `preview`<answer> holds exactly meta.answer_value (also the answer column of preview) on every row: score model output against that string.
The system prompt is ThinkMorph's VLM_THINK_SYSTEM_PROMPT from its inferencer.py, verbatim (GEN_THINK_SYSTEM_PROMPT there has the same text), including its leading and trailing newline. The markers are plain strings, not tokenizer special tokens; the prompt writes </image_end> and the data writes <image_end>, exactly as the ThinkMorph-7B checkpoint was trained.
preview shows the same rows with one column per slot: input_image_i for the inputs; for each of the K = num_steps rounds, the plan thought_j and its target target_image_j; and the read-back in thought_1 on every row.
meta holds the per-row sidecar: task, scene_id and geometry_uid (the scene and geometry keys; the split key is named in the split paragraph below), trajectory_id (a camera-path or sample label, empty where the source has none), num_steps, num_input_images, answer_type, answer_value, majority_class_rate, target_image_kind, target_px, est_tokens, licence, split (train / validation, the Hub split names), supervision_kind (full_interleaved / visual_aux / visual_only) and filter_flags. majority_class_rate is the share of the task's most frequent answer_value among its training rows: it measures answer skew and is not a guessing baseline (where a task mixes question types or each row has its own options it can be far below chance); compare scores with the text-only baselines below.
Per-row license in meta: cc-by-nc-4.0 10,854.
Flags on released rows (filter_flags in meta and preview, comma-separated):
Training with a BAGEL-family loader
Every row here has one input image (num_input_images is 1), so the stock ThinkMorph UnifiedEditIterableDataset (https://github.com/ThinkMorph/ThinkMorph: image_list[0] as input, image_list[j+1] after output_text_list[j]) and the IPT release's version (which reads num_input_images) both read it as intended. Mixed with a source whose rows have more than one input image, only a loader that reads num_input_images is correct.
The stock BAGEL edit loader (ByteDance-Seed/Bagel) cannot train these rows: it never reads output_text_list and expects each instruction_list element to be a list of paraphrases.
parquet_info.json keys each training chunk as <source>/<split>/<file>, here zebra_jigsaw/train/chunk_00000.parquet, with row-group counts read from the parquet footers. The loader matches a chunk only when its key equals the path it builds, os.path.join(data_dir, file), and skips a chunk with no key without a warning: a source that is alone in its group then fails with IndexError: list index out of range, and in a mixed group it adds no rows. Download into a directory named after the source, not after the repository:
from huggingface_hub import snapshot_download
snapshot_download("yrlyrl/spatial-mmcot-zebra_jigsaw", repo_type="dataset", local_dir="<root>/zebra_jigsaw",
allow_patterns=["train/*", "validation/*", "parquet_info.json", "reports/known_issues/*"])Then either run from <root> with data_dir: zebra_jigsaw/train and parquet_info_path: zebra_jigsaw/parquet_info.json, or rebuild the index with absolute keys and use an absolute data_dir:
import json, os
root = "/abs/path/to/root" # the directory that holds zebra_jigsaw/
info = json.load(open(os.path.join(root, "zebra_jigsaw", "parquet_info.json")))
info = {os.path.join(root, k): v for k, v in info.items()}
json.dump(info, open(os.path.join(root, "zebra_jigsaw", "parquet_info_abs.json"), "w"))
# data_dir = os.path.join(root, "zebra_jigsaw", "train") (spelled exactly so, no trailing slash)
# parquet_info_path = os.path.join(root, "zebra_jigsaw", "parquet_info_abs.json")The Hugging Face cache (.../snapshots/<hash>/train/) or a folder named spatial-mmcot-zebra_jigsaw matches no key.
num_used_data counts chunk files, not rows: the loader repeats this source's file list up to that number, lists every (file, row group) pair, and deals whole row groups out, floor(R / worldsize) to each rank and floor(that / numworkers) to each DataLoader worker. The remainder is never read. This source has 1 training chunk file holding 83 row groups of up to 128 rows, so keep num_used_data large, e.g. the 128 of ThinkMorph's interleaved_reasoning.yaml (upstream's example.yaml asks for more than GPUs x workers); every row group is then read. Set to 1 and alone in its group on 8 GPUs with 4 workers, it reads only 64 of the 83 row groups. In a run that mixes sources, give each source the same multiple of its own training chunk-file count, e.g. 128 per file (128 here): the file list is repeated up to num_used_data entries, so a flat 128 for every source would read a two-file source's rows half as often as a one-file source's.
How the rows were chosen
The rows refused before conversion are exactly the upstream rows whose trace carries four reasoning images: the type-1 puzzles, where three of the four candidate assemblies are deliberately wrong and the trace names the answer before any image is drawn. That rule, not the ids in build/dropped.jsonl, reproduces the refused set.
Every removed row has one line, with its reason, in reports/:
S0raw lines in build/dropped.jsonl were refused before a release row existed, so their row_uid field holds the converter's key for the upstream record, not a 16-hex row_uid; lines from later steps carry the row_uid the row had. No removed row appears in meta or preview. For this source the key is built from upstream fields that repeat across rows (for most sources a hash of the question text), so it is not unique: the 10,862 S7.multi_option_targets lines carry 9,862 distinct keys. Those rows are counted with their reason but cannot be traced to individual upstream rows.
S8.k_zero and S8.chain_broken name what happened to the row, not which check removed the image; the lines in this build do not record whether it was the size, transparency or copy check.
<details><summary>Per-step counters of the conversion</summary>
21,899 upstream rows were read; S0raw refused 10,862 before a row existed and passed 11,037 to the first step. S0 runs once more, last, on the final bytes. The reason for every refused, dropped or quarantined row is in the files above.
</details>
The train/validation split keeps rows sharing a geometry_uid in meta on one side, and the assignment is frozen (splits/ in the summary repository). geometry_uid is a hash of the upstream target, the complete source image, so puzzles cut from one image stay on one side; scene_id here is a per-puzzle hash of the input. S12 saw 10,962 rows under 10,954 keys. No validation input image has the content of a training input image. The S12 run did not record whether its pixel-level near-copy test ran for this source, so near-copies are not ruled out.
Answer-prior balancing (S13)
Each (task, split) group is checked separately. An answer is the answer value compared as lower-cased text without a trailing full stop, with 'farther' read as 'further' and 'nearer' as 'closer' (for multiple choice, the option text, not the letter; where the candidates are drawn in the image, as in zebrajigsaw and zebratetris, the answer is the letter itself). An answer is real when it holds at least 5 rows and 2% of the group; k is the number of real answers. Answer step: the target is max(30%, 1/k) when k >= 2, and max(30%, 1/d) over the d distinct answers when k = 1; a validation group uses the larger of its own target and its task's train target. A group is cut only when k >= 1 and its most common answer holds more than the target plus 5 percentage points; every answer is then capped at one common count, chosen so that none exceeds the target, and smaller answers keep all their rows. At the answer step, a group at or below that trigger, or with no real answer (k = 0), is left as it is, so its most common answer can hold up to the target plus 5 percentage points. A task whose train group has exactly two real answers is instead cut, in every split, so that its two largest answers have equal counts, with no trigger. Rank and label steps: then, in a group where every option value of every row is a number, the rank of the correct option among the sorted values, and after it, in a group where every trained answer is an option label, the label, are each capped by the same cut-and-trigger rule on their own counts (own target, validation included): capped, never evened out, so two labels are cut only when one exceeds 55%, and then only down to 50%. These steps can also cut groups the answer step left whole, including k = 0 groups, and can raise an answer's final share above its target; the run fails if a real answer ends above the target plus 5 percentage points. A train group of at least 20 rows in which one answer holds 90% or more fails the run. PET (exactcellspet) instead cuts each (question type x turn direction) cell to equal counts of its two answers; a PET cell that shows only one answer is removed.
S13 removed no row from this source.
Text-only baselines
Accuracy of guessers that never see an image. For each task the released training rows are split into two fixed halves by a hash of row_uid; each guesser is fitted on one half and scored once on the other (one held-out half, not cross-validation; eval rows below). The reference is chance (the mean of 1 / number of options) where every row is multiple choice, and otherwise the eval-half accuracy of always giving the answer most common in the fit half (when a task's top answers are nearly tied, this need not be the task's most common answer; the line after the table gives that answer's validation score). Accuracies are recounted from the stored rates and eval rows, so they are exact. A task is flagged when a text-only guesser beats its reference by more than 0.15 (for a free-form task, a guesser other than the most common answer). A flagged task can be partly answered from the text alone; an unflagged task passed only these probes, which do not prove the text carries no answer. Report scores on every task next to this baseline.
Guessers: keywords: the most common answer per set of spatial words in the question; last_mentioned: the option named last in the question body; letter_prior: the most common answer letter; majority: the answer most common in the fit half; option_prior: the option text that won most often when shown; template: the most common answer per question wording (numbers masked, object names kept).
Spot-check (S14)
Pending. The S14 rows are chosen and flagged S14.sampled_qa in meta and preview; the human pass over them has not been signed off yet.
Citation
Zebra-CoT is by Ang Li, Charles Wang, Deqing Fu, Kaiyu Yue, Zikui Cai, Wang Bill Zhu, Ollie Liu, Peng Guo, Willie Neiswanger, Furong Huang, Tom Goldstein and Micah Goldblum, released under CC BY-NC 4.0; this repository is a converted subset of it (the changes are listed under the licence line above). Its card asks users to cite:
@inproceedings{li2026zebracot,
title={Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning},
author={Ang Li and Charles Wang and Deqing Fu and Kaiyu Yue and Zikui Cai and Wang Bill Zhu and Ollie Liu and Peng Guo and Willie Neiswanger and Furong Huang and Tom Goldstein and Micah Goldblum},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=c6XIVI3TiQ}
}The images are ImageNet's; please also cite J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li and L. Fei-Fei, "ImageNet: A Large-Scale Hierarchical Image Database", CVPR 2009 (https://image-net.org/).
Provenance
The release files were written by our conversion code (the code repository is not public yet), scripts/convert/export.py at commit b3fa1b993ca9, from build zebra_jigsaw_r3. The build was made by scripts/convert/run_source.py from the same repository at commit 76c78bd817c2. S10, S12 and S13 ran before the export; reports/export_manifest.json pins every input the export read by SHA-1 (build_manifest_sha1, s10_keep_sha1, s12_assignments_sha1, s13_balanced_keep_sha1).
Every row removed between upstream and this release has one line, with its reason, in reports/: build/dropped.jsonl (rows refused before conversion or dropped by a conversion step); build/quarantine.jsonl (rows set aside by S4c because an automatic check could not match the read-back's conclusion to the label); s10_dropped.jsonl (duplicates removed by S10); s10_label_conflicts.jsonl (rows S10 withheld because another row asks the identical question, options in the same order, of the same images with a different answer); s13_dropped.jsonl (rows removed by answer-prior balancing). known_issues/ lists rows with a measured problem (see Known issues); reports/ also holds the build manifest (absolute paths cut to basenames) and counters, the S14 sample list (s14_sample.tsv: rowuid, task, split) and `exportmanifest.json. Part of [yrlyrl/spatial-mmcot`](https://huggingface.co/datasets/yrlyrl/spatial-mmcot).
