VR
Models
All models matching “VR”Datasets
All datasets matching “VR”Vript
🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo]
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc).… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript.Vript_Multilingual
🎬 Vript: A Video Is Worth Thousands of Words [Github Repo]
We construct another fine-grained video-text dataset with 19.1K annotated high-resolution UGC videos (~677k clips) in multiple languages to be the Vript_Multilingual.
New in Vript_Multilingual:
Multilingual: zh (60%), en (17%), de (15%), ja (6%), ko (2%), ru (<1%), es (<1%), pt (<1%), jv (<1%), fr (<1%), id (<1%), vi (<1%)
More diverse and fine-grained categories: 113 categories (please check vript_CN-V2_meta.json)… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript_Multilingual.VRSBench
VRSBench
VRSBench is a Versatile Vision-Language Benchmark for Remote Sensing Image Understanding. It consists of 29,614 remote sensing images with detailed captions, 52,472 object refers, and 3123,221 visual question-answer pairs. It facilitates the training and evaluation of vision-language models across a broad spectrum of remote sensing image understanding tasks.
Using datasets
from datasets import load_dataset
fw = load_dataset("xiang709/VRSBench"… See the full description on the dataset page: https://huggingface.co/datasets/xiang709/VRSBench.yumi-umi-session-20261008-000612
UMI raw capture — October 8, 2026
Task: Pick up the red block and put it on the blue block.
10 episodes, 6,024 frames, nominal 30 FPS, 640 × 480 wrist RGB video. Captured with D405 and T265; LeRobot v3 raw-observation schema. Original depth arrays and full-rate sensor logs are retained in extras/. Capture configuration and measurement semantics are in session.json.
This is an archival upload of the completed local session session-20261008-000612, without conversion or filtering.… See the full description on the dataset page: https://huggingface.co/datasets/vruga/yumi-umi-session-20261008-000612.VRBenchThis repository contains the dataset described in the paper VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos.
Project page: https://VRBench.github.io
Vript_Chinese
🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo]
We construct a fine-grained video-text dataset with 44.7K annotated high-resolution videos (~293k clips) in Chinese. The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript_Chinese.
