datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse
Vchitect Team1
1Shanghai Artificial Intelligence Laboratory
Paper |
Project Page |
Data Overview
The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.
It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.vbvr-latent-cache-832x832x33f-t2v-only
VBVR Latent Cache (832×832 × 33f, Wan2.2-TI2V-5B VAE + UMT5-XXL)
Pre-encoded latent cache for the
Video-Reason/VBVR-Dataset
geometric / logical reasoning video corpus, prepared for Equilibrium Matching
(EqM) post-training of Wan-AI/Wan2.2-TI2V-5B-Diffusers on AWS Trainium2.
This is a working cache, not a primary dataset. It exists to skip the
~5 s/sample VAE+T5 encode cost during training. The original videos +
prompts live in the upstream VBVR-Dataset repo.
Source →… See the full description on the dataset page: https://huggingface.co/datasets/Central-Cat/vbvr-latent-cache-832x832x33f-t2v-only.Wan2.1-T2V-1.3B_vidprom_81x480x832_40step_5cfg_5.0shift_4tvbench-wan22-t2v-dense93-xpu
VBench dense93 — Wan2.2 T2V videos (Intel XPU / vLLM-Omni)
93 text-to-video generations for the VBench "dense93" prompt set, produced
with Wan2.2-T2V-A14B (Diffusers) served by vLLM-Omni on Intel XPU
(oneAPI, 720x1280 @ 16fps, 81 frames, 40 denoise steps, guidance 4.0/3.0,
boundary 0.875, flow shift 5.0, seed 42, MXFP8 linear + cache-dit +
Sage V3 hybrid attention with SDPA fallback on blocks 33,34,38,39).
One video per prompt; file name = <prompt>-0.mp4
(prompt list:… See the full description on the dataset page: https://huggingface.co/datasets/Yi30/vbench-wan22-t2v-dense93-xpu.ERCBench-Wan2.2-T2V-A14B-400
Wan2.2 T2V A14B generations for EEC-Bench (400 episodes)
Official Wan2.2-T2V-A14B generations for all 400 episodes in the frozen EEC-Bench benchmark queue.
Specifications
Model: (MoE text-to-video)
Resolution: 832x480 (480P)
Frame count: 81 frames per shot (~5.06 seconds at 16 fps)
Sampling: UniPC solver, 40 steps, shift=12, guide_scale=[3.0, 4.0]
Precision: BF16 official conversion, unquantized
Format: Per-shot video clips under and concatenated episode video… See the full description on the dataset page: https://huggingface.co/datasets/BlueSourceJY/ERCBench-Wan2.2-T2V-A14B-400.EvalCrafter_T2V_Dataset
EvalCrafter Text-to-Video (ECTV) Dataset 🎥📊
Code · Project Page · Huggingface Leaderboard · Paper@ArXiv · Prompt list
Welcome to the ECTV dataset! This repository contains around 10000 videos generated by various methods using the Prompt list. These videos have been evaluated using the innovative EvalCrafter framework, which assesses generative models across visual, content, and motion qualities using 17 objective metrics and subjective user opinions.
Dataset Details 📚… See the full description on the dataset page: https://huggingface.co/datasets/RaphaelLiu/EvalCrafter_T2V_Dataset.Vchitect_T2V_DataVerse_256p_8fps_wdshttps://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse resampled to 256p. Intended for training https://github.com/NilanEkanayake/TiTok-Video
wan_t2v_distillation_datasetWan2.2-T2V-Activations-FP4t2v-sampled-videosViBeThis repository contains the data presented in ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models.
ViBe Annotations and Category Labels
Order of Model Annotations
The annotations for the models in metadata.csv are organized in the following order:
animatelcm
zeroscopeV2_XL
show1
mora
animatelightning
animatemotionadapter
magictime
zeroscopeV2_576w
ms1.7b
hotshotxl
Step-Video-T2V-EvalThis dataset contains the data of the paper Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.
Code: https://github.com/stepfun-ai/Step-Video-T2V
Project page: https://yuewen.cn/videos
ChinaOpen1k-T2V
ChinaOpen-1k T2V retrieval
Multilingual video retrieval built from
ChinaOpen-1k:
1,092 Bilibili videos with human-written Chinese captions and their
English translations. Both language configs share the same videos, so
the two subsets form a controlled comparison.
Built by scripts/data/chinaopen_retrieval/create_data.py in
mteb.
Please cite the original dataset:
@inproceedings{chen2023chinaopen,
title = {ChinaOpen: A Dataset for Open-world Multimodal Learning},
author =… See the full description on the dataset page: https://huggingface.co/datasets/shriyasudhakar/ChinaOpen1k-T2V.t2v-distill-benchmark
T2V Distillation Benchmark
A benchmark comparison of Text-to-Video distillation/acceleration models based on Wan2.1-14B.
Models Compared
Model
Steps
Source
FastVideo CausalWan2.2
8
FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers
Krea Realtime-Video
4
krea/krea-realtime-video
LightX2V CausVid
9
lightx2v/Wan2.1-T2V-14B-CausVid
NVlabs rCM 14B
4
worstcoder/rcm-Wan
Helios-Distilled
3 (pyramid 2-2-2)
BestWishYsh/Helios-Distilled… See the full description on the dataset page: https://huggingface.co/datasets/hffordata/t2v-distill-benchmark.wan2.1_t2v_14b_2k
Wan2.1 Text-to-Video 14B 数据集 (2K样本)
数据集描述
这是一个包含 2,000 个 文本到视频生成对的数据集,来自 Wan2.1 T2V 14B 模型的前2000条数据。
数据集内容
视频文件: 2,000 个 MP4 格式视频
元数据文件: 2,000 个 JSON 格式元数据文件
总大小: 约 3.8 GB(视频) + 2.5 MB(元数据)
文件结构
data/
├── videos/ # 视频文件 (video_00000.mp4 到 video_01999.mp4)
└── metadata/ # 元数据文件 (video_00000.json 到 video_01999.json)
元数据结构
每个 JSON 文件包含以下字段:
{
"text_prompt": "文本提示词",
"video_id": "视频ID",
"generation_params": {… See the full description on the dataset page: https://huggingface.co/datasets/NZC415/wan2.1_t2v_14b_2k.T2V-CompBench-VideosFLARE-1k-Audio-T2VAFLARE-1k-Unified-T2VAT2Vcinepile-t2v-split_scenes_single_shot_uniform
Important Columns for Captioning
Caption_t2v_style: Expressive and long caption generated by Gemini Flash 2.5 for the extracted shot.
Caption_t2v_style_short: Short caption generated by Gemini Flash 2.5 for the extracted shot.
Avg-Aesthetic-Score-Laion-Aesthetics: Average (over frames) aesthetic score of the extracted shot from Laion Aesthetics.
Frame-Aesthetic-Scores-Laion-Aesthetics: Aesthetic scores of each frame of the extracted shot from Laion Aesthetics.… See the full description on the dataset page: https://huggingface.co/datasets/CinematicT2vData/cinepile-t2v-split_scenes_single_shot_uniform.azm-archive-20260909-t2v-needlabel-videos
t2v_data_needlabel_videos.tar
Backup of an existing dataset archive, preserving its original bytes.
File: t2v_data_needlabel_videos.tar
Size: 4,675,983,360 bytes
SHA256: 8672488cef23b55e2a70d3279327a38ddbbd8e2809661600d7c707d59b9e1e18
Verify the downloaded archive with sha256sum -c SHA256SUMS.
gfvc-vace-cogvideox-t2v
gfvc-vace-cogvideox-t2v
Mirror of the exact bmcore v24 holdout subset. Source: hugging4chang/gfvc-vace-synthetic-t2v-data.
All 2,500 local basenames match the public text-to-video subset. Public card explicitly excludes TalkVid/YouTube originals and real-person image-to-video. Preserve model terms; no blanket relicensing.
Contains 2500 synthetic video files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed:… See the full description on the dataset page: https://huggingface.co/datasets/34data/gfvc-vace-cogvideox-t2v.gfvc-vace-wan2-1-t2v
gfvc-vace-wan2-1-t2v
Mirror of the exact bmcore v24 holdout subset. Source: hugging4chang/gfvc-vace-synthetic-t2v-data.
All 2,500 local basenames match the public text-to-video subset. Public card explicitly excludes TalkVid/YouTube originals and real-person image-to-video. Preserve model terms; no blanket relicensing.
Contains 2500 synthetic video files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed:… See the full description on the dataset page: https://huggingface.co/datasets/34data/gfvc-vace-wan2-1-t2v.t2v_data_v2
DenseDPO T2V Broad-Pair Dataset (v2)
Cross-model text-to-video (T2V) generation pairs for training video reward
models (RM) and DPO-style preference learning.
The HF Dataset Viewer renders each row as prompt + two videos side-by-side.
Generation task
All videos are generated T2V from a shared text prompt. For every pair,
both videos share the same prompt, so the primary comparison axis is the model
identity itself.
Plan-A tier structure
Models are grouped into… See the full description on the dataset page: https://huggingface.co/datasets/qgfvadfuvads/t2v_data_v2.t2v-batch300Wan_2.2_T2V_10steps_GGUFmixkit_filtered_6k_wan1.3_t2vThe raw mixkit dataset is downloaded and preprocessed following instructions from https://github.com/tianweiy/CausVid/tree/master
Info_Wan_Video_2.2_T2V-A14B
Model Index by Creator
423748 Page
Model
Base Model
Full Model Page
Archive Link
wan2.2,t2v,low,zzzyixuan.
Wan Video 2.2 T2V-A14B
View
View
Version Links
Model
Version
Base Model
Version Link
wan2.2,t2v,low,zzzyixuan.
v1.0
Wan Video 2.2 T2V-A14B
View
Aaron_PP Page
Model
Base Model
Full Model Page
Archive Link
NSFW WAN 2.2 T2V Bunny girl, red patent leather tights, black high stockings, red high heels
Wan Video 2.2… See the full description on the dataset page: https://huggingface.co/datasets/ApacheOne/Info_Wan_Video_2.2_T2V-A14B.mixkit-t2vWan21_CausVid_14B_T2V_lora_rank32.safetensors
