Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gasstation /gs-videos-v3text10K<n<100K1 likes9.7k downloads5mo agoHugging Face02facebook /PE-Video PE Video Dataset (PVD) [📃 Tech Report] [📂 Github] The PE Video Dataset (PVD) is a large-scale collection of 1 million diverse videos, featuring 120,000+ expertly annotated clips. The dataset was introduced in our paper "Perception Encoder". Overview PE Video Dataset (PVD) comprises 1M high quality and diverse videos. Among them, 120K videos are accompanied by automated and human-verified annotations. and all videos are accompanied with video description and keywords.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PE-Video.text100K<n<1M52 likes6.2k downloads1y agoHugging Face03teowu /LSVQ-videosThis is an unofficial copy of the videos in the LSVQ dataset (Ying et al, CVPR, 2021), the largest dataset available for Non-reference Video Quality Assessment (NR-VQA); this is to facilitate research studies on this dataset given that we have received several reports that the original links of the dataset is not available anymore. See FAST-VQA (Wu et al, ECCV, 2022) or DOVER (Wu et al, ICCV, 2023) repo on its converted labels (i.e. quality scores for videos). The file links to the labels in… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LSVQ-videos.textvideo-classification10K<n<100K5 likes1.9k downloads3y agoHugging Face04LongVideo-Reason /longvideo_eval_videos Long-RL: Scaling RL to Long Sequences (Evaluation Dataset - for research only) Data Distribution We strategically construct a high-quality dataset with CoT annotations for long video reasoning, named LongVideo-Reason. Leveraging a powerful VLM (NVILA-8B) and a leading open-source reasoning LLM, we develop a dataset comprising 52K high-quality Question-Reasoning-Answer pairs for long videos. We use 18K high-quality samples for Long-CoT-SFT to initialize… See the full description on the dataset page: https://huggingface.co/datasets/LongVideo-Reason/longvideo_eval_videos.text1K<n<10K1 likes1.5k downloads1y agoHugging Face05gasstation /gs-videos-v2text10K<n<100K0 likes1.5k downloads8mo agoHugging Face06ShareGPTVideo /train_raw_video ShareGPTVideo Raw ActivityNet Videos for Train data All dataset and models can be found at ShareGPTVideo. Contents: Due to our scene split, we provide our processed activityNet videos corresponding to test frames in train video frames the processing script is process_activitynet.py textquestion-answering10K<n<100K2 likes1.2k downloads2y agoHugging Face07aslhlf /Videos_FL_0216text10K<n<100K0 likes868 downloads8mo agoHugging Face08gasstation /generated-videostext100K<n<1M0 likes758 downloads11mo agoHugging Face09BAAI-DataCube /Physics-aware-videos中文 README 🧱 Physics-aware Video Dataset This dataset is a high-quality real-world video dataset focused on physical phenomena, designed for learning and evaluating physical laws from videos. It primarily covers the following classic physical processes: 🧱 Rigid-body motion / 🌊 Fluid dynamics 💫 Collision and rebound 💥 Explosion and burst phenomena 🌫️ Smoke, dust, and particle scattering 🌍 Gravity and inertia (falling, rolling, acceleration, etc.) The dataset is constructed… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/Physics-aware-videos.text10K<n<100K5 likes538 downloads10mo agoHugging Face10rrustlee /videomme_1fps_336_336textn<1K0 likes510 downloads2y agoHugging Face11ApolloVideo /videomindtext100K<n<1M0 likes476 downloads7mo agoHugging Face12Kwai-Keye /VideoTemp-o3 VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos Illustration of the agentic pipeline in VideoTemp-o3. Given a video QA pair, the model performs on-demand temporal grounding to locate the most relevant segment, then refines it iteratively. Finally, it produces a reliable answer grounded in the pertinent visual evidence. Data Source The question and answer pairs used for training VideoTemp-o3 are sourced from… See the full description on the dataset page: https://huggingface.co/datasets/Kwai-Keye/VideoTemp-o3.text10K<n<100K1 likes448 downloads5mo agoHugging Face13aslhlf /Videos_FL_0204text1K<n<10K0 likes400 downloads8mo agoHugging Face14VisionXLab /FIRM-Video FIRM-Video-SFT-90K This repository releases the 90K SFT data for FIRM-Video. The dataset covers three key evaluation dimensions: Instruction Following (IF): whether the generated video accurately follows the text prompt. Visual Quality (VQ): perceptual and technical quality, including clarity, sharpness, artifacts, flicker, and overall visual fidelity. World Coherence (WC): whether the video is coherent with commonsense, temporal consistency, physical plausibility, and… See the full description on the dataset page: https://huggingface.co/datasets/VisionXLab/FIRM-Video.image1M<n<10M0 likes344 downloads2mo agoHugging Face15aslhlf /Videos_FL_0225text10K<n<100K0 likes265 downloads8mo agoHugging Face16aslhlf /Videos_FL_0221text10K<n<100K0 likes189 downloads8mo agoHugging Face17HIT-TMG /VideoVista-CoTs VideoVista-CoTs This repository contains VideoVista-CoTs, used in Uni-MoE-2.0 training. This dataset samples a portion of data from LLaVA-Video-178K, SEED-Bench-R1, SR-91K, and STAR, and uses our automatic Video QA generation framework to perform multi-step reasoning annotations for filtered complex questions. The automatic video QA generation codes and our VideoVista series are presented in VideoVista Family Citation If you find VideoVista-CulturalLingo useful for your… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/VideoVista-CoTs.textvideo-text-to-text10K<n<100K1 likes139 downloads10mo agoHugging Face18findcard12138 /open_video_datatext100K<n<1M0 likes136 downloads2y agoHugging Face19Darknsu /mead_hdtf_400_merge_video_audio_frames_onlyimage1M<n<10M0 likes132 downloads4mo agoHugging Face20BAAI-DataCube /Hand-action-videos中文 README ✋ Hand Action Video Dataset ✋ Hand Action Video Dataset This dataset is a real-world video dataset focused on hand actions (Hand Action Video Dataset), with an emphasis on common operations in hand–object interaction scenarios. It contains a large number of videos featuring: ✋ Clean and clear hand actions 🧱 Explicit hand–object interaction relationships The dataset is constructed via video retrieval on **Datacube ** combined with automatic filtering using multimodal… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/Hand-action-videos.text100K<n<1M0 likes129 downloads10mo agoHugging Face21ShareGPTVideo /test_raw_video_data ShareGPTVideo Raw Videos for Testing data All dataset and models can be found at ShareGPTVideo. Contents: In case of need, this contains raw videos corresponding to test frames in Test video frames textquestion-answering1K<n<10K2 likes115 downloads2y agoHugging Face22bitmind /aislop-videostext1K<n<10K0 likes115 downloads1y agoHugging Face23mukul54 /video_chatgpt_activitynet_videostext1K<n<10K0 likes113 downloads2y agoHugging Face24Tinghong-Ye66 /dance_videotext100K<n<1M0 likes98 downloads5mo agoHugging Face25aidealab /aidealab-videojp-eval AIdeaLab VideoJP 評価再現用データ はじめに このリポジトリはAIdeaLab VideoJPのFVDを測定するためのデータを 集めました。再現手順を次のとおりに示します。 評価方法 まず、評価用ライブラリをダウンロードします。 git clone https://github.com/JunyaoHu/common_metrics_on_video_quality ダウンロードできたら、ライブラリのインストール手順を踏んで、インストールします。 インストールしたら、同じディレクトリに次のファイルをコピーしてください evaluate_videos.py videos.tar gen_ja.tar コピーできたら、videos.tarとgen_ja.tarを展開します。 tar xf videos.tar tar xf gen_ja.tar 最後にevaluate_videos.pyを実行すると、FVDが表示されるはずです。 おまけ: 評価用映像の作り方… See the full description on the dataset page: https://huggingface.co/datasets/aidealab/aidealab-videojp-eval.texttext-to-video1K<n<10K0 likes88 downloads1y agoHugging Face26dhyun22 /video-diffusion-perceptionimage10M<n<100M4 likes78 downloads6mo agoHugging Face27ZhangAo /video-ctext1K<n<10K0 likes76 downloads2y agoHugging Face28mxxxxxxxxxxxxxxxxx /galaxy_video_clip_featurestext1M<n<10M0 likes73 downloads1y agoHugging Face29Yiming1234 /VoT-video-latent-archiveimage100K<n<1M0 likes70 downloads6mo agoHugging Face30thisnick /nsfw-video-still-caption-grid-onlyimage10K<n<100K14 likes65 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.