CrossVideoReasoning/SYNCR
SYNCR SYNCR is a simulator-grounded framework for cross-video reasoning: questions that cannot be answered from any single video, but require aligning events, matching identities, comparing motion, or integrating partial observations across several videos. Because the videos are produced in simulation, every answer is derived from environment state rather than from human annotation. The same generators produce both an evaluation benchmark and a training set over disjoint videos… See the full description on the dataset page: https://huggingface.co/datasets/CrossVideoReasoning/SYNCR.
SYNCR
SYNCR is a simulator-grounded framework for cross-video reasoning: questions that cannot be answered from any single video, but require aligning events, matching identities, comparing motion, or integrating partial observations across several videos.
Because the videos are produced in simulation, every answer is derived from environment state rather than from human annotation. The same generators produce both an evaluation benchmark and a training set over disjoint videos, so supervision can be studied without contaminating evaluation.
Contents
The evaluation set draws on 4,827 distinct video files and the training set on 14,956, with no video shared between the two splits. (These are unions over tasks; per-task counts sum to more, because some videos are used by more than one task.)
Loading
The benchmark and the training set have different schemas, so they are published as two configs:
from datasets import load_dataset
test = load_dataset("CrossVideoReasoning/SYNCR", split="test") # 4,000 questions
train = load_dataset("CrossVideoReasoning/SYNCR", "training", split="train") # 15,960 examplesGetting the videos
Video paths in the annotations are relative to the repository root and resolve under videos/:
videos/clvr/video_validation/video_11194.mp4
videos/clvr/video_train/video_05825.mp4
videos/habitat_data/route_plan/...
videos/kubric_data/scene_144/cam2.mp41. Habitat and Kubric — extract the archives in this repository:
mkdir -p videos && for f in *_data.tar; do tar -xf "$f" -C videos/; done2. CLEVRER — these clips are not redistributed here. Download them from the official CLEVRER release at http://clevrer.csail.mit.edu and place them so the paths above resolve:
mkdir -p videos/clvr
unzip video_validation.zip -d videos/clvr/ # 1,987 clips used by the benchmark
unzip video_train.zip -d videos/clvr/ # 5,879 clips used by the training setWhich sources each split needs:
Tasks
Eight tasks in four reasoning families, 500 evaluation questions each.
Numerical Comparison presents five options; every other task presents four. Correct answers are balanced across option positions (975 each for A–D, plus 100 at E for the five-option task).
Schema
Each row of test/metadata.jsonl:
Video order is significant. videos[0] is "Video 1" in the question text, and so on.
Evaluation protocol
Models answer zero-shot with the videos supplied in order and labelled Video 1:, Video 2:, … We score the parsed final option with deterministic decoding where supported.
Two notes that materially affect measured accuracy:
- Allow enough generation budget. Some models emit long chains of thought before the answer and, if truncated, produce no parseable option — which scores as wrong and understates the model. This is most pronounced on Synchronization. A budget of 8,192 new tokens is sufficient for the models we tested; 2,048 is not.
- Score unparseable answers explicitly. Report them separately rather than silently counting them as incorrect, particularly for base models.
Training set
train/sft_all8_answer_only.json is a chat-formatted mixture over all eight tasks (15,960 examples), with targets containing only the answer line. Each message list interleaves text and video content blocks; the video field carries the same relative paths as the benchmark.
Video sources and licensing
Videos are rendered from three simulators, each with its own upstream terms:
- CLEVRER — Yi et al., ICLR 2020 — http://clevrer.csail.mit.edu
- Kubric — Greff et al., CVPR 2022
- Habitat — Savva et al., ICCV 2019 / Szot et al., NeurIPS 2021
The annotations in this repository are released under CC BY 4.0. CLEVRER clips are not redistributed here and remain under their original terms. Rendered Habitat and Kubric clips remain subject to the licenses of those simulators and their scene assets; consult those before redistribution.
Changes from the previous release
This release supersedes the earlier version of this dataset and is not a drop-in replacement:
- The evaluation set is now 4,000 questions (500 per task), replacing the previous 8,163-example single split, and a 15,960-example training split has been added.
- Tasks were renamed and regenerated: CLEVRER collision comparison → Numerical Comparison, Kubric spatial tracking → Spatial Measurement, Habitat agent/object tracking → Object Re-identification.
- Video paths are now repository-relative rather than absolute.
Results reported in the accompanying paper correspond to this version.
