fengyee/camera-motion-dataset-and-benchmark
Camera Motion Dataset and Benchmark This repository packages the dataset and benchmark release for our paper "Geometry-Guided Camera Motion Understanding in VideoLLMs", accepted to the CVPR 2026 Workshop on Pixel-level Video Understanding in the Wild (PVUW2026), as a Hugging Face dataset repo. Derived from the KlingTeam/MultiCamVideo-Dataset. It contains: WebDataset tar shards for train, val, and test Split manifests (train.json, val.json, test.json) that map source clips to… See the full description on the dataset page: https://huggingface.co/datasets/fengyee/camera-motion-dataset-and-benchmark.
Camera Motion Dataset and Benchmark
This repository packages the dataset and benchmark release for our paper ["Geometry-Guided Camera Motion Understanding in VideoLLMs"](https://arxiv.org/abs/2603.13119), accepted to the [CVPR 2026 Workshop on Pixel-level Video Understanding in the Wild (PVUW2026)](https://pvuw.github.io/), as a Hugging Face dataset repo. Derived from the KlingTeam/MultiCamVideo-Dataset.
It contains:
- WebDataset tar shards for
train,val, andtest - Split manifests (
train.json,val.json,test.json) that map source clips to sample ids - A multiple-choice VQA benchmark (
cam-motion.jsonl) - Small example scripts under
examples/
Files
Dataset Format
Each WebDataset sample stores 8 RGB frames resized to 560 x 560 and a JSON metadata record:
<sample_id>.jpg-00
<sample_id>.jpg-01
...
<sample_id>.jpg-07
<sample_id>.jsonThe JSON sidecar has the form:
{
"label": "dolly in, truck right",
"id": "00ae5c50b2b54704ad833996c20c055d"
}The manifests provide the lookup back to source provenance:
{
"video_path": "path/to/source/video.mp4",
"clip_index": 2,
"label": "dolly in, truck right",
"id": "00ae5c50b2b54704ad833996c20c055d"
}cam-motion.jsonl contains multiple-choice benchmark rows like:
{
"video": "path/to/source/video.mp4",
"clip_index": 2,
"conversations": [
{
"from": "human",
"value": "<video>\nIdentify the camera motion depicted in the video using standard cinematographic terminology.\nOptions:\n(A) ...\n(B) ...\n(C) ...\n(D) ..."
},
{
"from": "gpt",
"value": "Answer: (A) ..."
}
]
}The video field in the benchmark is a provenance string copied from dataset construction. After release, it should be matched against video_path in the manifest files rather than treated as a valid local path.
Current Split Layout
- Manifest rows:
train=9819,val=1227,test=1228(12274 total) - Materialized shard samples: matches manifests exactly (0 missing)
- Benchmark rows:
12274(one row per unique clip)
All splits are deduplicated and non-overlapping.
The example and validation scripts in this repo show how to inspect these counts locally before publishing.
Usage
Download the dataset repo locally, then point webdataset at the shard files:
import io
import json
from pathlib import Path
from PIL import Image
import webdataset as wds
dataset_root = Path("camera-motion-dataset-and-benchmark")
shards = sorted(str(p) for p in dataset_root.glob("train-*.tar"))
def decode_sample(sample):
frames = [
Image.open(io.BytesIO(sample[f"jpg-{i:02d}"])).convert("RGB")
for i in range(8)
]
meta = json.loads(sample["json"])
return {
"id": meta["id"],
"label": meta["label"],
"frames": frames,
}
dataset = wds.WebDataset(shards, shardshuffle=False).map(decode_sample)
sample = next(iter(dataset))
print(sample["id"], sample["label"], len(sample["frames"]))To inspect the benchmark:
python examples/read_benchmark.py --dataset-root . --show 3To inspect one tar shard:
python examples/inspect_tar.py --tar train-000000.tarSuggested Hugging Face Workflow
- Stage a clean dataset repo from the source project using the local helper scripts.
- If you are rebuilding corrected manifests first, point
stage_release.pyat that dataset directory with--dataset-dir. - Create a private dataset repo first.
- Upload the staged folder.
- Inspect the dataset card, files, and example scripts on the Hub.
- Make the repo public only after the release audit looks correct.
Citation
@article{feng2026geometry,
title = {Geometry-Guided Camera Motion Understanding in VideoLLMs},
author = {Feng, Haoan and Musunuri, Sri Harsha and Su, Guan-Ming},
journal = {arXiv preprint arXiv:2603.13119},
year = {2026},
url = {https://arxiv.org/abs/2603.13119}
}