Kangverse/Sekai2_Real_World
Sekai2 Real World This repository releases the reproducible URL/timestamp metadata and paired camera-pose/caption annotations for the perspective-video portion of Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling. Resources: ๐ Project Page ยท ๐ป GitHub ยท ๐ Paper The perspective MP4 clips are not redistributed here. Each row in sekai2_clips.csv provides the source URL and the exact half-open frame range [start_frame, end_frame) in a canonical 30โฆ See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.
Sekai2 Real World
This repository releases the reproducible URL/timestamp metadata and paired camera-pose/caption annotations for the perspective-video portion of Sekai2. See the paper: [Sekai2: From World Exploration to Interactive World Modeling](https://arxiv.org/abs/2608.09449).
Resources: ๐ Project Page ยท ๐ป GitHub ยท ๐ Paper
The perspective MP4 clips are not redistributed here. Each row in sekai2_clips.csv provides the source URL and the exact half-open frame range [start_frame, end_frame) in a canonical 30 fps time base. The accompanying script downloads each source video and reconstructs the clip used by the released pose and caption.
Where the three release components live
The static supplement is not exposed as a separate dataset label. Its nonduplicate clips are merged into sekai2. The 5,597 retained static-supplement clip IDs are listed in metadata/sekai2_static_clips.csv, so users can remove that subset with a simple ID filter.
Metadata
sekai2_clips.csv: 127,884 YouTube-reconstructable perspective clips fromsekaiandsekai2.sekai2_dataset.csv: complete public metadata for all released Sekai2 data, includingsekai,sekai2, and the 20-hour panoramic subset.metadata/sekai2_static_clips.csv: IDs of the optional static supplement for the current v2 release.metadata/annotation_shards.csv: size and SHA-256 of every active v2 annotation tar.metadata/v1/: immutable shard checksums and static-clip IDs for the legacyannotations/release.metadata/v2/: shard checksums and static-clip IDs for the currentannotations-v2/release. The two CSVs directly undermetadata/mirror these v2 files for backward compatibility.
Core reconstruction columns include clip_id, url, fps, start_frame, end_frame, start_time, end_time, and num_frames. Annotation lookup uses the relative pose_path and caption_path columns. For perspective clips these point to public Hugging Face tar members (annotations-v2/...tar::pose/...); for panoramic clips they point to ModelScope files.
Each annotation tar stores paired members:
pose/<clip_stem>.npz
caption/<clip_stem>.jsonAnnotation versions
annotations/is the legacy v1 release. Its pose NPZ files do not include camera intrinsics.annotations-v2/is the current release. Its pose NPZ files include both per-sample camera intrinsics (intrinsics) and camera-to-world transforms (cam_c2w). The root CSV manifests currently point to this v2 release.
Each released sekai2 pose NPZ uses the Sekai camera format:
intrinsics # (T, 3, 3), float32
cam_c2w # (T, 4, 4), float32T is the number of camera-trajectory samples and is clip-dependent; it is not padded or truncated to a fixed length. The two arrays always have the same T. The sekai2 annotation directory is split into 17 tar files with up to 5,000 paired clips per shard.
The previous annotations/ tree is retained temporarily as unreferenced legacy data. Current CSV paths and metadata/annotation_shards.csv identify the active annotations-v2/ tree.
For sekai, every released pose comes from the 47,699-clip pose_vipe_smoothed set and every caption comes from caption_k25. For the public caption copies, annotation_meta.model and annotation_meta.machine_id are removed; all semantic caption fields are kept.
Reconstruction convention
For every perspective row:
start_time = start_frame / 30
end_time = end_frame / 30
num_frames = end_frame - start_frameDo not use stream-copy cutting: GOP/keyframe seeking can shift clip boundaries. Use the provided reconstruction script, which performs frame-accurate seeking, normalizes to a 30 fps output time base at 1280ร720, forces the declared frame count, and then verifies the result. A full PyAV audit found that every perspective release clip has a 30 fps video stream base rate. Some YouTube source videos were originally higher-fps or variable-rate; the release metadata therefore uses the audited 30 fps output clip length, not the raw source frame rate. The panoramic subset is distributed as full videos and is not represented as URL/timestamp cuts.
Citation
Please cite the Sekai2 paper when using this release:
@article{he2026sekai2,
title = {Sekai2: From World Exploration to Interactive World Modeling},
author = {He, Kang and Peng, Wenshuo and Gao, Zihui and Tan, Jiaming and
Zhang, Kaipeng and Ge, Yongtao},
journal = {arXiv preprint arXiv:2608.09449},
year = {2026}
}