datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/ali-vilab/VACE-Benchmark.waymo_vace_train_causal_subset10vace-aug-224-gr00t
VACE-augmented RoboCasa dataset — 224 episodes (GR00T / LeRobot v2.1)
The exact 224-episode training set used to fine-tune GR00T-1.5 in the VACE-augmentation
baseline. Built by swapping a target object into RoboCasa pick-and-place episodes with the
VACE video-diffusion model, then converting to the GR00T "gr00t_views" (LeRobot v2.1) format.
224 episodes, 62,445 frames, 3 camera views (left_view / right_view / wrist_view
= RoboCasa robot0_agentview_left / robot0_agentview_right… See the full description on the dataset page: https://huggingface.co/datasets/mlnha/vace-aug-224-gr00t.tfix-vacevace-aug
vace-aug — VACE object-swap augmentation for RoboCasa
Code and packed inputs for swapping a target mesh into RoboCasa pick-and-place episodes
with the VACE video-diffusion model (8 target objects x 32 episodes x 3 cameras), built as
a baseline to compare against MimicGen-based augmentation.
aug32/ the augmentation pipeline — start at aug32/README.md
VACE/ ali-vilab/VACE fork carrying the batch flags this needs… See the full description on the dataset page: https://huggingface.co/datasets/mlnha/vace-aug.Wan2.1-VACEVACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/Manding0/VACE-Benchmark.gfvc-vace-wan2-1-t2v
gfvc-vace-wan2-1-t2v
Mirror of the exact bmcore v24 holdout subset. Source: hugging4chang/gfvc-vace-synthetic-t2v-data.
All 2,500 local basenames match the public text-to-video subset. Public card explicitly excludes TalkVid/YouTube originals and real-person image-to-video. Preserve model terms; no blanket relicensing.
Contains 2500 synthetic video files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed:… See the full description on the dataset page: https://huggingface.co/datasets/34data/gfvc-vace-wan2-1-t2v.gfvc-vace-cogvideox-t2v
gfvc-vace-cogvideox-t2v
Mirror of the exact bmcore v24 holdout subset. Source: hugging4chang/gfvc-vace-synthetic-t2v-data.
All 2,500 local basenames match the public text-to-video subset. Public card explicitly excludes TalkVid/YouTube originals and real-person image-to-video. Preserve model terms; no blanket relicensing.
Contains 2500 synthetic video files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed:… See the full description on the dataset page: https://huggingface.co/datasets/34data/gfvc-vace-cogvideox-t2v.VACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/BonnyWang/VACE-Benchmark.Scenery_Anime_Bright_VACE_Depth_V2V_Captioned_50_samples
transformed from https://huggingface.co/datasets/svjack/Scenery_Shot_Videos_Captioned_50_samples
VACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/zywer/VACE-Benchmark.waymo_vace_train_causal_subset5waymo_vace_val_causal_50gfvc-vace-synthetic-t2v-data
GFVC VACE Synthetic T2V Data
This public repository contains 5,000 text-to-video synthetic clips only. It
does not contain TalkVid/YouTube source videos or image-to-video clips
initialized from real-person reference images.
Contents
Collection
Generator family
Videos
Bytes
wan2_1_t2v
Wan2.1 T2V
2500
6422035594
cogvideox_t2v
CogVideoX T2V
2500
353365051
data_manifest.csv records media identity and probe metadata. Its
source_path values are… See the full description on the dataset page: https://huggingface.co/datasets/hugging4chang/gfvc-vace-synthetic-t2v-data.VACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/gzhxiaobai/VACE-Benchmark.Xiang_IceCream_Kontext_OmniConsistency_VACE_Videos
Reference on
transform by: https://huggingface.co/svjack/Kontext_OmniConsistency_lora
use:
"transform it into 3D Chibi style",
"transform it into American Cartoon style",
"transform it into Chinese Ink style",
"transform it into Clay Toy style",
"transform it into Fabric style",
"transform it into Ghibli style",
"transform it into Irasutoya style",
"transform it into Jojo style","transform it into LEGO style",
"transform it into Line style",
"transform it into Macaron style",
"transform… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Xiang_IceCream_Kontext_OmniConsistency_VACE_Videos.scene-extrapolation-dl3dv-vace14b
Scene Extrapolation: DL3DV → Wan2.1-VACE-14B (warp + inpaint)
Generate novel-view video that extends a real scene beyond the observed region. Observed frames are
depth-warped (DepthAnythingV2 + COLMAP poses) into target viewpoints → a control video with holes +
an inpaint mask → Wan2.1-VACE-14B generates the unobserved region.
Method
observed frames → forward-splat warp into target poses → [control w/ holes | mask] → VACE-14B inpaints
Data: 1,500 DL3DV scenes… See the full description on the dataset page: https://huggingface.co/datasets/luuuulinnnn/scene-extrapolation-dl3dv-vace14b.ckpts_vaceABot-PhysWorld_eval_640_480_HYPIR_vace_skeleton_3250stepCut_Fruit_VACE_Depth_V2V_Captioned
waymo_vace_train_causal_fullskriptwan_vace_transformation_videosvace_combinemask_3250stepVACE_A2-BenchXiang_Float_After_Tomorrow_Head_RMBG_SPLITED_VACE_OutPainting_videoswaymo_vace_val_causal_fullXiang_Float_After_Tomorrow_Head_RMBG_SPLITED_VACE_OutPainting_srcsXiang_Float_After_Tomorrow_Mix_SPLITED_VACE_OutPainting_videos
