datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.VidProM
Summary
This is the dataset proposed in our paper VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models (NeurIPS 2024).
VidProM is the first dataset featuring 1.67 million unique text-to-video prompts and 6.69 million videos generated from 4 different state-of-the-art diffusion models.
It inspires many exciting new research areas, such as Text-to-Video Prompt Engineering, Efficient Video Generation, Fake Video Detection, and Video Copy… See the full description on the dataset page: https://huggingface.co/datasets/WenhaoWang/VidProM.StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.trex-visualizer
T-Rex Dataset Visualizer
A browseable subset of the T-Rex dataset — Tactile-Rich Bimanual Dexterous
Manipulation — collected on a bimanual Dexmate Vega-1 robot equipped with
two Sharpa Wave dexterous hands.
This visualizer subset contains 3,838 short trajectory clips drawn from the
full 100-hour T-Rex collection, organized by (verb, object, hand) so you can
quickly inspect coverage across motion primitives and object categories.
For the full dataset (multi-view RGB, robot… See the full description on the dataset page: https://huggingface.co/datasets/Beakerman0101/trex-visualizer.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.VPData
VideoPainter
This repository contains the implementation of the paper "VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control"
Keywords: Video Inpainting, Video Editing, Video Generation
Yuxuan Bian12, Zhaoyang Zhang1‡, Xuan Ju2, Mingdeng Cao3, Liangbin Xie4, Ying Shan1, Qiang Xu2✉
1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong 3The University of Tokyo 4University of Macau ‡Project Lead ✉Corresponding Author
Your… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/VPData.UltraVideo
UltraVideo: High-Quality UHD 4K Video Dataset
🤓 Project | 📑 Paper | 🤗 Hugging Face (UltraVideo Dataset)) | 🤗 Hugging Face (UltraVideo-Long Dataset)) | 🤗 Hugging Face (UltraWan-1K/4K Weights)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
🎋 Click below image to watch the 4K demo video.
🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo.Sparkle
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou
📦 Dataset
Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper.
The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.MIntRec
Dataset details
In real-world conversational interactions, we usually combine information from multiple modalities (e.g., text, video, audio) to help analyze human intentions. Though intent analysis has been widely explored in the Natural Language Processing community, there is a scarcity of data for multimodal intent analysis. Thus, we provide a novel multimodal intent benchmark dataset, MIntRec, to boom the research. To the best of our knowledge, it is the first multimodal intent… See the full description on the dataset page: https://huggingface.co/datasets/THU-IAR/MIntRec.IntPhys2
IntPhys 2
Dataset |
Hugging Face |
Paper |
Blog
IntPhys 2 is a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2… See the full description on the dataset page: https://huggingface.co/datasets/facebook/IntPhys2.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.SVBench
Dataset Card for SVBench
This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository.
Dataset Details
Dataset Description
SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.FilmBench
FilmBench — Video Generation Benchmark Dataset
📢 Update (2026-08-02)
Added English prompt files: filmbench_prompts_en.csv is now available with English prompts.
filmbench_prompts_en.csv (1,169 rows): Prompt-level table with English prompts. Columns: uid, task, movie_type (English), en_prompt, reference_url.
📢 Update (2026-07-31)
Fixed a batch of misaligned prompts in filmbench_videos.csv: the zh_prompt column has been recalibrated against the… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/FilmBench.Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Truebones-ZOO-Annotations.ReactiveGWM-Datasets
ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models
📚 Datasets-Introduction
ReactiveGWM-Datasets is the strategy-aligned training corpus that powers
ReactiveGWM, a game world
model that decouples player control from NPC autonomy. To learn that
decoupling, the model needs supervision that pairs each gameplay clip with
both a per-frame action stream (what the player did) and a high-level
NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.wlasl_reduced
Reduced WLASL Dataset
This dataset is a task-specific reduced version of the WLASL dataset,
constructed for American Sign Language (ASL) recognition experiments.
Contents
videos/Video clips organized by gloss label.
metadata.csvPer-sample metadata including:
file path
gloss label
fps (after normalization, if applied)
video resolution
normalized bounding box coordinates
gloss_map.jsonMapping from gloss labels to integer class IDs.
Dataset Construction… See the full description on the dataset page: https://huggingface.co/datasets/jherng/wlasl_reduced.processed_vnhnVideoArtifactDetectionVideo Artifact Detection dataset using source videos from LongVideoBench. Video level labels are given in labels.csv (training) and labels_test.csv (testing). Localized artifact regions for burst artifacts are given in the artifact_ranges column.
Note that labels.csv contains additional source videos from LongVideoBench that are not included in this repository. You may visit the LongVideoBench page for the additional videos.
For additional questions, please email palmerla@usc.edu.
VideoDRUltraVideo-Long
UltraVideo: High-Quality UHD 4K Video Dataset
🤓 Project | 📑 Paper | 🤗 Hugging Face (UltraVideo Dataset)) | 🤗 Hugging Face (UltraVideo-Long Dataset)) | 🤗 Hugging Face (UltraWan-1K/4K Weights)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
🎋 Click below image to watch the 4K demo video.
🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo-Long.ActivityForensics
[CVPR 2026] ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
Author: Peijun Bao, Anwei Luo, Gang Pan, Alex C. Kot, Xudong Jiang
[Project Page]
[Paper]
[Supp]
[Code]
[Dataset]
If this dataset is useful for your work, please consider citing our paper
@inproceedings{bao2026activityforensics,
title={ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos},
author={Bao, Peijun and Luo… See the full description on the dataset page: https://huggingface.co/datasets/ActivityForensics/ActivityForensics.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.FullBenchLSV
LSV: LabSuperVision Benchmark
Dataset Description
LSV is a multi-view video dataset of wet-lab biology experiments, captured from both first-person (XMglass smart glasses) and third-person (DJI action camera) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors.
The dataset is designed for research on:
Protocol compliance… See the full description on the dataset page: https://huggingface.co/datasets/YinkaiW/LSV.KABR
Dataset Card for KABR: In-Situ Dataset for Kenyan Animal Behavior Recognition from Drone Videos
Dataset Summary
We present a novel high-quality dataset for animal behavior recognition from drone videos.
The dataset is focused on Kenyan wildlife and contains behaviors of giraffes, plains zebras, and Grevy's zebras.
The dataset consists of more than 10 hours of annotated videos, and it includes eight different classes, encompassing seven types of animal behavior and an… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR.Sekai2_Real_World
Sekai2 Real World
This repository releases the reproducible URL/timestamp metadata and paired
camera-pose/caption annotations for the perspective-video portion of
Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling.
Resources: 🌐 Project Page · 💻 GitHub · 📄 Paper
The perspective MP4 clips are not redistributed here. Each row in
sekai2_clips.csv provides the source URL and the exact half-open frame range
[start_frame, end_frame) in a canonical 30… See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.GameplayQA
GameplayQA: A Decision-Dense POV-Synced Multi-Video
Understanding Benchmark of 3D Virtual Agents
Yunzhe Wang
Runhui Xu
Kexin Zheng
Tianyi Zhang
Jayavibhav N. Kogundi
Soham Hans
Volkan Ustun
University of Southern California
ACL 2026
Corresponding Author: yunzhewa@usc.edu
Overview
GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.
