Sta8is/cityscapes_seq_video
Cityscapes Sequence Video Short video clips built from the Cityscapes sequence data, packaged for training video generation / world models. Contents cityscapes/ ├── train/ │ ├── videos/ 2975 clips (cs_sec_*.mp4) │ ├── metas/ general_prompt.txt — text caption shared by all clips │ └── t5_xxl/ general_prompt.pickle — T5-XXL embedding of that caption └── val/ ├── videos/ 500 clips (cs_sec_*.mp4) ├── metas/… See the full description on the dataset page: https://huggingface.co/datasets/Sta8is/cityscapes_seq_video.
Cityscapes Sequence Video
Short video clips built from the Cityscapes sequence data, packaged for training video generation / world models.
Contents
cityscapes/
├── train/
│ ├── videos/ 2975 clips (cs_sec_*.mp4)
│ ├── metas/ general_prompt.txt — text caption shared by all clips
│ └── t5_xxl/ general_prompt.pickle — T5-XXL embedding of that caption
└── val/
├── videos/ 500 clips (cs_sec_*.mp4)
├── metas/ general_prompt.txt
└── t5_xxl/ general_prompt.pickle- Clips: MP4, 2048×1024, 30 frames @ 16 fps (~1.9 s each).
- Caption:
"The video is captured from a camera mounted on a car. The camera is facing forward." - t5_xxl:
pickleholding a list with onefloat16array of shape(21, 1024)— the T5-XXL encoding of the caption, reusable as a conditioning tensor.
License
Videos are derived from the Cityscapes dataset and remain subject to the Cityscapes license. Non-commercial / research use only; see the original terms before redistributing.
