Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acvlab /ABot-World-Explorer-500h ABot World Explorer 500h ABot World Explorer 500h contains 30,969 action-conditioned video episodes associated with the data infrastructure described in ABot-World-0. Each episode preserves an MP4, dataset-native keyboard actions, captions, and one COLMAP text sparse model. Dataset facts Item Value Episodes 30,969 Source objects 185,814 Semantic splits None License Apache-2.0 The repository name is an identifier, not an audited… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h.text10K<n<100K29 likes60k downloads2mo agoHugging Face02OraRL /OraRL-Data OraRL-Data [🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code] We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL. It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.imagevideo-text-to-text10K<n<100K1 likes21k downloads1mo agoHugging Face03MCG-NJU /VideoChat3-LV116k VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments. The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.textvideo-text-to-text1K<n<10K15 likes19k downloads4d agoHugging Face04Mutonix /Vript 🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo] We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc).… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript.textvideo-classification100K<n<1M27 likes17k downloads2y agoHugging Face05Lijiaxin0111 /M3_VOS [CVPR 2025] M3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation If you like our project, please give us a star ⭐ on GitHub for the latest update. 💡 Description Venue: CVPR2025 Repository: 🛠️Tool, 🏠Page Paper: arxiv.org/html/2412.13803v2 Point of Contact: Jiaxin Li , Zixuan Chen 📁 Structure This dataset contains annotated videos and images for object segmentation tasks with phase transition information. The directory… See the full description on the dataset page: https://huggingface.co/datasets/Lijiaxin0111/M3_VOS.imagevideo-classificationn<1K1 likes16k downloads10mo agoHugging Face06ShareGPT4Video /ShareGPT4Video ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.imagevisual-question-answering10K<n<100K204 likes10k downloads2y agoHugging Face07prolongvid /ProLongVid_data Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only allow the use of this dataset for academic research and education purpose. Paper: For more details, please check our paper Code: For training recipe and other update, please refer to github repo. Citation @inproceedings{ wang2025prolongvid, title={ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning}, author={Rui Wang… See the full description on the dataset page: https://huggingface.co/datasets/prolongvid/ProLongVid_data.text1M<n<10M1 likes10k downloads11mo agoHugging Face08yihongs /VOST-TAS [NeurIPS 2025] Tracking and Understanding Object Transformations If you like our project, please give us a star ⭐ on GitHub for the latest update. 💡 Description Dataset Visualizations: GitHub Paper: arXiv:2511.04678 Project Page: tubelet-graph.github.io Project Repository: GitHub Point of Contact: Yihong Sun 📊 Dataset Overview VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.imagevideo-classificationn<1K0 likes9.9k downloads9mo agoHugging Face09markov-ai /gaming-500-hours Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play session, trimmed to pure gameplay — login screens, launchers, desktop, collection-app references, and any watching/streaming are removed. In-game menus, lobbies, loading, and cutscenes are retained as part of the session. Workflows: 776 Total gameplay: 494.7 hours Distinct games: 168 Clip duration (min): median 24.0, p90 90.9, max 457.7 Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.tabularn<1K259 likes8.8k downloads3mo agoHugging Face10Tianli /robovqatext100K<n<1M5 likes7.7k downloads1y agoHugging Face11InternRobotics /RoboInter-Data RoboInter-Data: Intermediate Representation Annotations for Robot Manipulation Rich, dense, per-frame intermediate representation annotations for robot manipulation, built on top of DROID and RH20T. Developed as part of the RoboInter project. You can try our Online demo. The annotations cover 230k episodes and include: subtasks, primitive skills, segmentation, gripper/object bounding boxes, placement proposals, affordance boxes, grasp poses, traces, contact points, etc. And each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/RoboInter-Data.textrobotics1K<n<10K16 likes7.7k downloads8mo agoHugging Face12CG-Bench /CG-Benchgated CG-Bench Project Website: https://cg-bench.github.io/leaderboard/GitHub Repository: https://github.com/CG-Bench/CG-Bench (includes running code) Summary We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question answering in long videos, addressing the limitations of existing benchmarks that focus primarily on short videos and rely on multiple-choice questions (MCQs). These limitations allow models to answer by elimination rather than genuine… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-Bench.tabularvisual-question-answering10K<n<100K10 likes7.7k downloads2y agoHugging Face13AIGVDBench /AIGVDBenchtext10K<n<100K9 likes6.9k downloads6mo agoHugging Face14mohantesting /video-quality-scored Image-to-Video Quality-Scored Clips A collection of prompted image-to-video samples with quality-evaluation metadata. Each sample pairs a first frame (the I2V conditioning image) with one or both of: a generated video produced by a video model from the first frame + prompt an original clip (the reference/source video the prompt was authored around) A subset of the samples also carry per-clip quality scores: an overall quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.imagetext-to-video1K<n<10K0 likes6.6k downloads4mo agoHugging Face15juyil /AVQA-videos AVQA — Audio-Visual Question Answering (videos + annotations) A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life audio-visual question answering over short in-the-wild clips. The original release ships only the QA annotations and expects users to collect the source videos from VGGSound themselves. This repository bundles the source video clips together with the official train/val annotations, so the dataset is usable without any YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.tabularvisual-question-answering10K<n<100K1 likes6k downloads4mo agoHugging Face16wlyu /AnyCam_MCV_480p AnyCam_MCV_480p Object-centric camera-control training pairs built from the MultiCamVideo dataset (ReCamMaster; 13,600 Unreal Engine 5 scenes, each filmed by 10 synchronised cameras with a person animating in the foreground). Each sample is a source/target pair of two different cameras of the same scene, plus one control video for each side. The control videos show a single anchor object: on the source side it is cut out of the source video with its SAM 2 mask, on the target… See the full description on the dataset page: https://huggingface.co/datasets/wlyu/AnyCam_MCV_480p.tabulartext-to-video1K<n<10K0 likes5.6k downloads1d agoHugging Face17artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K0 likes5.5k downloads2mo agoHugging Face18MIN-Lab /minWM-datatext1K<n<10K1 likes5.4k downloads5mo agoHugging Face19Harland /OmniVChat OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue &nbsp;&nbsp;&nbsp;&nbsp; 1 The Chinese University of Hong Kong 2 Alibaba Token Hub, Alibaba Group 3 Shanghai Jiao Tong University 4 Shanghai Innovation Institute 5 Zhejiang University OmniVChat (Omni Video Chat) is the task of native audio-visual dialogue: an omni model directly and simultaneously receives audio and video from a user and returns text. The user's query is inside the audio and… See the full description on the dataset page: https://huggingface.co/datasets/Harland/OmniVChat.textvideo-text-to-text1K<n<10K56 likes4.8k downloads15d agoHugging Face20UnFaZeD07 /Music-AVQAtabular10K<n<100K0 likes4.6k downloads8mo agoHugging Face21VLM2Vec /MSR-VTTClone from "friedrichor/MSR-VTT". MSRVTT contains 10K video clips and 200K captions. We adopt the standard 1K-A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text-Video Retrieval field. Train: train_7k: 7,010 videos, 140,200 captions train_9k: 9,000 videos, 180,000 captions Test: test_1k: 1,000 videos, 1,000 captions 🌟 Citation @inproceedings{xu2016msrvtt, title={Msr-vtt: A large video description dataset… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSR-VTT.tabulartext-to-video10K<n<100K4 likes4.4k downloads1y agoHugging Face22Vividbot /vivid-video-instructtext10K<n<100K2 likes4.4k downloads2y agoHugging Face23facebook /wearable-aigated EgoWearBench Dataset (ECCV 2026) Part of the Wearable AI Workshop at ECCV 2026. A benchmark of egocentric (first-person, head-mounted wearable camera) videos paired with three complementary video question-answering tasks for evaluating wearable-AI assistants on real-world everyday activity videos. ▶ Baseline code & evaluation scripts: see starter_kit/README.md. The starter kit ships inside this repo, so git clone gives you the code and the data together. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/wearable-ai.textvisual-question-answering1K<n<10K27 likes3.7k downloads28d agoHugging Face24VLM2Vec /DiDeMoClone from friedrichor/DiDeMo. About DiDeMo contains 10K long-form videos from Flickr. For each video, ~4 short sentences are annotated in temporal order. We follow the existing works to concatenate those short sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 8,395 videos, 8,395 captions (concatenate from 33,005 short captions) Val: 1,065 videos, 1,065 captions (concatenate from 4,290 short captions) (We don't have… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/DiDeMo.texttext-to-video1K<n<10K0 likes3.4k downloads1y agoHugging Face25buaaplay /SVCBench SVCBench: Streaming Video Counting Benchmark This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline. Project Page: https://buaa-colalab.github.io/SVCBench/ Code: https://github.com/buaa-colalab/SVCBench Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/buaaplay/SVCBench.tabularvideo-classification1K<n<10K10 likes3.4k downloads3mo agoHugging Face26yale-nlp /MMVU MMVU: Measuring Expert-Level Multi-Discipline Video Understanding 🌐 Homepage • 🥇 Leaderboard • 📖 Paper • 🤗 Data 📰 News 2025-01-21: We are excited to release the MMVU paper, dataset, and evaluation code! 👋 Overview Why MMVU Benchmark? Despite the rapid progress of foundation models in both text-based and image-based expert reasoning, there is a clear gap in evaluating these models’ capabilities in specialized-domain video understanding.… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/MMVU.textvideo-text-to-text1K<n<10K59 likes3.4k downloads2y agoHugging Face27hyf015 /EgoExoLearnNOTE: Videos in huggingface are unprocessed, full-size videos. For benchmark and gaze alignment, we use processed 25fps videos. For processed data and code for benchmark, please visit the github page. EgoExoLearn This repository contains the video data of the following paper: EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong… See the full description on the dataset page: https://huggingface.co/datasets/hyf015/EgoExoLearn.tabularvideo-classificationn<1K7 likes3.2k downloads2y agoHugging Face28davidwdw /fa-task63-candidate18-raw-rollouts-320487ddbbc8 task63-candidate18-raw-rollouts Versioned fleet archive. Canonical recipe: datasets/task63_all_v1/5f4427151e8a Tier: raw Codegen collection rollouts, train instances 1-300 (candidate18): per-attempt actions, logs and evaluation outputs plus runner logs; producer of task63_all_v1/success_v1 and recorded source of the dev20 replay gate Use the exact recorded revision and verify SHA256SUMS. This package is a snapshot, not a live directory mirror. tabularn<1K0 likes3.1k downloads6d agoHugging Face29davidwdw /fa-task01-v14-train300-dense-raw-episodes-13c62ec5cf12 task01-v14-train300-dense-raw-episodes Versioned fleet archive. Canonical recipe: datasets/task01-v14-train300-dense/collection-20260923 Tier: raw Codegen collection episodes, 300 instances / 305 attempts (success, failure and infrastructure-error attempts): per-attempt policy actions/events, evaluator logs, evaluation json and per-attempt LeRobot export Use the exact recorded revision and verify SHA256SUMS. This package is a snapshot, not a live directory mirror. tabularn<1K0 likes2.9k downloads6d agoHugging Face30VLM2Vec /MSVDClone from "friedrichor/MSVD". MSVD contains 1,970 videos, each of which is paired with ~40 captions. We adopt the official split: Train: 1,200 videos, 48,774 captions Val: 100 videos, 4,290 captions Test: 670 videos, 27,763 captions 🌟 Citation @inproceedings{chen2011collecting, title={Collecting highly parallel data for paraphrase evaluation}, author={Chen, David and Dolan, William B}, booktitle={Proceedings of the Annual Meeting of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSVD.texttext-to-video1K<n<10K4 likes2.8k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.