Team Ai
20 results

VR

Mutonix /Vript 🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo] We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc).… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript.textvideo-classification100K<n<1M27 likes15k downloads2y agoHugging FaceMutonix /Vript_Multilingual 🎬 Vript: A Video Is Worth Thousands of Words [Github Repo] We construct another fine-grained video-text dataset with 19.1K annotated high-resolution UGC videos (~677k clips) in multiple languages to be the Vript_Multilingual. New in Vript_Multilingual: Multilingual: zh (60%), en (17%), de (15%), ja (6%), ko (2%), ru (<1%), es (<1%), pt (<1%), jv (<1%), fr (<1%), id (<1%), vi (<1%) More diverse and fine-grained categories: 113 categories (please check vript_CN-V2_meta.json)… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript_Multilingual.textvideo-classification100K<n<1M7 likes7.3k downloads2y agoHugging Facexiang709 /VRSBench VRSBench VRSBench is a Versatile Vision-Language Benchmark for Remote Sensing Image Understanding. It consists of 29,614 remote sensing images with detailed captions, 52,472 object refers, and 3123,221 visual question-answer pairs. It facilitates the training and evaluation of vision-language models across a broad spectrum of remote sensing image understanding tasks. Using datasets from datasets import load_dataset fw = load_dataset("xiang709/VRSBench"… See the full description on the dataset page: https://huggingface.co/datasets/xiang709/VRSBench.visual-question-answering10K<n<100K19 likes3.4k downloads4mo agoHugging Facevruga /yumi-umi-session-20261008-000612 UMI raw capture — October 8, 2026 Task: Pick up the red block and put it on the blue block. 10 episodes, 6,024 frames, nominal 30 FPS, 640 × 480 wrist RGB video. Captured with D405 and T265; LeRobot v3 raw-observation schema. Original depth arrays and full-rate sensor logs are retained in extras/. Capture configuration and measurement semantics are in session.json. This is an archival upload of the completed local session session-20261008-000612, without conversion or filtering.… See the full description on the dataset page: https://huggingface.co/datasets/vruga/yumi-umi-session-20261008-000612.tabular1K<n<10K0 likes3.4k downloads3d agoHugging FaceOpenGVLab /VRBenchThis repository contains the dataset described in the paper VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos. Project page: https://VRBench.github.io video-text-to-text5 likes3.1k downloads1y agoHugging FaceMutonix /Vript_Chinese 🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo] We construct a fine-grained video-text dataset with 44.7K annotated high-resolution videos (~293k clips) in Chinese. The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript_Chinese.textvideo-classification100K<n<1M16 likes2.7k downloads2y agoHugging Face