t2v
Datasets
All datasets matching “t2v”Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse
Vchitect Team1
1Shanghai Artificial Intelligence Laboratory
Paper |
Project Page |
Data Overview
The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.
It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.vbvr-latent-cache-832x832x33f-t2v-only
VBVR Latent Cache (832×832 × 33f, Wan2.2-TI2V-5B VAE + UMT5-XXL)
Pre-encoded latent cache for the
Video-Reason/VBVR-Dataset
geometric / logical reasoning video corpus, prepared for Equilibrium Matching
(EqM) post-training of Wan-AI/Wan2.2-TI2V-5B-Diffusers on AWS Trainium2.
This is a working cache, not a primary dataset. It exists to skip the
~5 s/sample VAE+T5 encode cost during training. The original videos +
prompts live in the upstream VBVR-Dataset repo.
Source →… See the full description on the dataset page: https://huggingface.co/datasets/Central-Cat/vbvr-latent-cache-832x832x33f-t2v-only.Wan2.1-T2V-1.3B_vidprom_81x480x832_40step_5cfg_5.0shift_4tvbench-wan22-t2v-dense93-xpu
VBench dense93 — Wan2.2 T2V videos (Intel XPU / vLLM-Omni)
93 text-to-video generations for the VBench "dense93" prompt set, produced
with Wan2.2-T2V-A14B (Diffusers) served by vLLM-Omni on Intel XPU
(oneAPI, 720x1280 @ 16fps, 81 frames, 40 denoise steps, guidance 4.0/3.0,
boundary 0.875, flow shift 5.0, seed 42, MXFP8 linear + cache-dit +
Sage V3 hybrid attention with SDPA fallback on blocks 33,34,38,39).
One video per prompt; file name = <prompt>-0.mp4
(prompt list:… See the full description on the dataset page: https://huggingface.co/datasets/Yi30/vbench-wan22-t2v-dense93-xpu.ERCBench-Wan2.2-T2V-A14B-400
Wan2.2 T2V A14B generations for EEC-Bench (400 episodes)
Official Wan2.2-T2V-A14B generations for all 400 episodes in the frozen EEC-Bench benchmark queue.
Specifications
Model: (MoE text-to-video)
Resolution: 832x480 (480P)
Frame count: 81 frames per shot (~5.06 seconds at 16 fps)
Sampling: UniPC solver, 40 steps, shift=12, guide_scale=[3.0, 4.0]
Precision: BF16 official conversion, unquantized
Format: Per-shot video clips under and concatenated episode video… See the full description on the dataset page: https://huggingface.co/datasets/BlueSourceJY/ERCBench-Wan2.2-T2V-A14B-400.EvalCrafter_T2V_Dataset
EvalCrafter Text-to-Video (ECTV) Dataset 🎥📊
Code · Project Page · Huggingface Leaderboard · Paper@ArXiv · Prompt list
Welcome to the ECTV dataset! This repository contains around 10000 videos generated by various methods using the Prompt list. These videos have been evaluated using the innovative EvalCrafter framework, which assesses generative models across visual, content, and motion qualities using 17 objective metrics and subjective user opinions.
Dataset Details 📚… See the full description on the dataset page: https://huggingface.co/datasets/RaphaelLiu/EvalCrafter_T2V_Dataset.
