Team Ai
Datasetpublic

Sta8is/cityscapes_seq_video

Cityscapes Sequence Video Short video clips built from the Cityscapes sequence data, packaged for training video generation / world models. Contents cityscapes/ ├── train/ │ ├── videos/ 2975 clips (cs_sec_*.mp4) │ ├── metas/ general_prompt.txt — text caption shared by all clips │ └── t5_xxl/ general_prompt.pickle — T5-XXL embedding of that caption └── val/ ├── videos/ 500 clips (cs_sec_*.mp4) ├── metas/… See the full description on the dataset page: https://huggingface.co/datasets/Sta8is/cityscapes_seq_video.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes955downloads
Dataset Card

Cityscapes Sequence Video

Short video clips built from the Cityscapes sequence data, packaged for training video generation / world models.

Contents

cityscapes/
├── train/
│   ├── videos/      2975 clips  (cs_sec_*.mp4)
│   ├── metas/       general_prompt.txt   — text caption shared by all clips
│   └── t5_xxl/      general_prompt.pickle — T5-XXL embedding of that caption
└── val/
    ├── videos/      500 clips  (cs_sec_*.mp4)
    ├── metas/       general_prompt.txt
    └── t5_xxl/      general_prompt.pickle
  • —Clips: MP4, 2048×1024, 30 frames @ 16 fps (~1.9 s each).
  • —Caption: "The video is captured from a camera mounted on a car. The camera is facing forward."
  • —t5_xxl: pickle holding a list with one float16 array of shape (21, 1024) — the T5-XXL encoding of the caption, reusable as a conditioning tensor.

License

Videos are derived from the Cityscapes dataset and remain subject to the Cityscapes license. Non-commercial / research use only; see the original terms before redistributing.