Team Ai
Datasetpublic

Kangverse/Sekai2_Real_World

Sekai2 Real World This repository releases the reproducible URL/timestamp metadata and paired camera-pose/caption annotations for the perspective-video portion of Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling. Resources: ๐ŸŒ Project Page ยท ๐Ÿ’ป GitHub ยท ๐Ÿ“„ Paper The perspective MP4 clips are not redistributed here. Each row in sekai2_clips.csv provides the source URL and the exact half-open frame range [start_frame, end_frame) in a canonical 30โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes1.3kdownloads
Dataset Card

Sekai2 Real World

This repository releases the reproducible URL/timestamp metadata and paired camera-pose/caption annotations for the perspective-video portion of Sekai2. See the paper: [Sekai2: From World Exploration to Interactive World Modeling](https://arxiv.org/abs/2608.09449).

Resources: ๐ŸŒ Project Page ยท ๐Ÿ’ป GitHub ยท ๐Ÿ“„ Paper

The perspective MP4 clips are not redistributed here. Each row in sekai2_clips.csv provides the source URL and the exact half-open frame range [start_frame, end_frame) in a canonical 30 fps time base. The accompanying script downloads each source video and reconstructs the clip used by the released pose and caption.

Where the three release components live

ComponentClipsLocation
sekai47,699This repository: annotations-v2/sekai/
sekai2 (new crawl + static supplement)80,185This repository: annotations-v2/sekai2/
20-hour panoramic subset173ModelScope: kangverse/Sekai2_Panoramics

The static supplement is not exposed as a separate dataset label. Its nonduplicate clips are merged into sekai2. The 5,597 retained static-supplement clip IDs are listed in metadata/sekai2_static_clips.csv, so users can remove that subset with a simple ID filter.

Metadata

  • โ€”sekai2_clips.csv: 127,884 YouTube-reconstructable perspective clips from sekai and sekai2.
  • โ€”sekai2_dataset.csv: complete public metadata for all released Sekai2 data, including sekai, sekai2, and the 20-hour panoramic subset.
  • โ€”metadata/sekai2_static_clips.csv: IDs of the optional static supplement for the current v2 release.
  • โ€”metadata/annotation_shards.csv: size and SHA-256 of every active v2 annotation tar.
  • โ€”metadata/v1/: immutable shard checksums and static-clip IDs for the legacy annotations/ release.
  • โ€”metadata/v2/: shard checksums and static-clip IDs for the current annotations-v2/ release. The two CSVs directly under metadata/ mirror these v2 files for backward compatibility.

Core reconstruction columns include clip_id, url, fps, start_frame, end_frame, start_time, end_time, and num_frames. Annotation lookup uses the relative pose_path and caption_path columns. For perspective clips these point to public Hugging Face tar members (annotations-v2/...tar::pose/...); for panoramic clips they point to ModelScope files.

Each annotation tar stores paired members:

text
pose/<clip_stem>.npz
caption/<clip_stem>.json

Annotation versions

  • โ€”annotations/ is the legacy v1 release. Its pose NPZ files do not include camera intrinsics.
  • โ€”annotations-v2/ is the current release. Its pose NPZ files include both per-sample camera intrinsics (intrinsics) and camera-to-world transforms (cam_c2w). The root CSV manifests currently point to this v2 release.

Each released sekai2 pose NPZ uses the Sekai camera format:

python
intrinsics  # (T, 3, 3), float32
cam_c2w     # (T, 4, 4), float32

T is the number of camera-trajectory samples and is clip-dependent; it is not padded or truncated to a fixed length. The two arrays always have the same T. The sekai2 annotation directory is split into 17 tar files with up to 5,000 paired clips per shard.

The previous annotations/ tree is retained temporarily as unreferenced legacy data. Current CSV paths and metadata/annotation_shards.csv identify the active annotations-v2/ tree.

For sekai, every released pose comes from the 47,699-clip pose_vipe_smoothed set and every caption comes from caption_k25. For the public caption copies, annotation_meta.model and annotation_meta.machine_id are removed; all semantic caption fields are kept.

Reconstruction convention

For every perspective row:

text
start_time = start_frame / 30
end_time   = end_frame / 30
num_frames = end_frame - start_frame

Do not use stream-copy cutting: GOP/keyframe seeking can shift clip boundaries. Use the provided reconstruction script, which performs frame-accurate seeking, normalizes to a 30 fps output time base at 1280ร—720, forces the declared frame count, and then verifies the result. A full PyAV audit found that every perspective release clip has a 30 fps video stream base rate. Some YouTube source videos were originally higher-fps or variable-rate; the release metadata therefore uses the audited 30 fps output clip length, not the raw source frame rate. The panoramic subset is distributed as full videos and is not represented as URL/timestamp cuts.

Citation

Please cite the Sekai2 paper when using this release:

bibtex
@article{he2026sekai2,
  title   = {Sekai2: From World Exploration to Interactive World Modeling},
  author  = {He, Kang and Peng, Wenshuo and Gao, Zihui and Tan, Jiaming and
             Zhang, Kaipeng and Ge, Yongtao},
  journal = {arXiv preprint arXiv:2608.09449},
  year    = {2026}
}