Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stanford-vision-lab /gpicgated GPIC: A Giant Permissive Image Corpus for Visual Generation Keshigeyan&nbsp;Chandrasegaran*1,&nbsp; Kyle&nbsp;Sargent*1,&nbsp; Suchir&nbsp;Agarwal1,&nbsp; Michael&nbsp;Jang1,&nbsp; Michael&nbsp;Poli1,2,&nbsp; Juan&nbsp;Carlos&nbsp;Niebles1,4,&nbsp; Justin&nbsp;Johnson3,&nbsp; Jiajun&nbsp;Wu1,&nbsp; Li&nbsp;Fei-Fei1 1&nbsp;Stanford University&nbsp;&nbsp; 2&nbsp;Radical Numerics&nbsp;&nbsp; 3&nbsp;University of Michigan&nbsp;&nbsp; 4&nbsp;Salesforce… See the full description on the dataset page: https://huggingface.co/datasets/stanford-vision-lab/gpic.162 likes165k downloads2mo agoHugging Face02hf-vision /course-assetsimagen<1K9 likes130k downloads2y agoHugging Face03CCTV-Vision-Team /pipeline-cctv-analytics2 likes29k downloads3mo agoHugging Face04nyu-visionx /Cambrian-10M Cambrian-10M Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-10M is a comprehensive dataset designed for instruction tuning, particularly in multimodal settings involving visual interaction data. The dataset is crafted to address the scarcity of high-quality multimodal instruction-tuning data and to maintain the language abilities of multimodal large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-10M.visual-question-answering1M<n<10M131 likes21k downloads2y agoHugging Face05LeonardoBenitez /VisionUnlearningEvaluationTestbeds1 likes17k downloads2mo agoHugging Face06horde-research /kaz-vision-50kimage0 likes14k downloads1y agoHugging Face07sensenova /SenseNova-Vision-Corpus-50M Vision as Unified Multimodal Generation English | 简体中文 This repository contains the dataset for the paper Vision as Unified Multimodal Generation. SenseNova Vision Corpus 50M Overview SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-Vision-Corpus-50M.imageany-to-anyn<1K58 likes14k downloads1mo agoHugging Face08ut-vision /EgoBraingated [ICLR 2026] EgoBrain: Synergizing Minds and Eyes For Human Action Understanding 東京大学 The Univerisity of Tokyo X 微軟亞洲研究院 Microsoft Research Asia Nie Lin · Yansen Wang · Dongqi Han · Weibang Jiang · Jingyuan Li · Ryosuke Furuta · Yoichi Sato* · Dongsheng Li* · *(Co-corresponding authors)* This is the official dataset repository of our ICLR 2026 paper "EgoBrain: Synergizing Minds and Eyes For Human Action… See the full description on the dataset page: https://huggingface.co/datasets/ut-vision/EgoBrain.videon<1K13 likes12k downloads3mo agoHugging Face09nyu-visionx /VSI-Bench Dataset arXiv Website Code VSI-Bench VSI-Bench-Debiased v1 [!IMPORTANT] [Aug. 9, 2026] PROVENANCE UPDATE: The existing "Debiased" subset is VSI-Bench-Debiased v1, a designer-in-the-loop manual pilot created with bespoke per-question-type filtering heuristics. It predates and was not generated by the automated Iterative Bias Pruning (IBP) algorithm. We retain v1 for reproducibility and will version any future automated subset separately.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-Bench.textvisual-question-answering10K<n<100K71 likes8.6k downloads2mo agoHugging Face10nyu-visionx /CV-Bench Cambrian Vision-Centric Benchmark (CV-Bench) This repository contains the Cambrian Vision-Centric Benchmark (CV-Bench), introduced in Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. Files The test*.parquet files contain the dataset annotations and images pre-loaded for processing with HF Datasets. These can be loaded in 3 different configurations using… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/CV-Bench.imagevisual-question-answering1K<n<10K48 likes7.7k downloads1y agoHugging Face11nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes6k downloads2y agoHugging Face12lmarena-ai /VisionArena-Chat VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes 200k single and multi-turn chats between users and VLM's collected on Chatbot Arena. WARNING: Images may contain inappropriate content. Dataset Details 200K conversations 45 VLM's 138 languages ~43k unique images Question Category Tags (Captioning, OCR, Entity Recognition, Coding, Homework, Diagram, Humor, Creative Writing, Refusal) Dataset Description 200,000… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/VisionArena-Chat.imagevisual-question-answering100K<n<1M15 likes5.1k downloads2y agoHugging Face13Asklv /OpenMath-Vision-CoT-10kimage10K<n<100K1 likes5k downloads9mo agoHugging Face14nyu-visionx /VSI-590K VSI-590K Website | Paper | GitHub | Models Authors: Shusheng Yang*, Jihan Yang*, Pinzhi Huang†, Ellis Brown†, et al. VSI-590K is a large-scale spatially-focused instruction-tuning dataset focusing on spatial reasoning. The dataset is curated from diverse sources and carefully annotated. Quick Start import json # Load from JSONL file with open('vsi_590k.jsonl', 'r') as f: for line in f: sample = json.loads(line.strip()) print(sample) break… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-590K.visual-question-answering100K<n<1M26 likes4.5k downloads11mo agoHugging Face15yuanqianhao /Vision-OPD-6K Vision-OPD-6K: Training Data for Vision-OPD Overview Vision-OPD proposes a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy, without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use. Vision-OPD instantiates two conditional policies from the same MLLM: A crop-conditioned teacher that observes the evidence-centered crop as a privileged… See the full description on the dataset page: https://huggingface.co/datasets/yuanqianhao/Vision-OPD-6K.text1K<n<10K11 likes2.7k downloads4mo agoHugging Face16nyu-visionx /Cambrian-S-3M Cambrian-S-3M TLDR: This is a collection of open-source video instruction tuning data used in Cambrian-S's third training stage. Overview Cambrian-S-3M combines three video instruction datasets: Cambrian-S-3M LLaVA-Video-178K LLaVA-Hound (ShareGPTVideo) Prerequisites Hugging Face CLI: pip install -U "huggingface_hub[cli]==0.36.0" Sufficient disk space (~5 TB recommended) hf command should be available after installing huggingface_hub Setup… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-S-3M.7 likes2.6k downloads9mo agoHugging Face17metomorfinov /clyde2-visiongeospatial0 likes2.3k downloads1mo agoHugging Face18ServiceNow /ui-vision UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction Introduction Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/ui-vision.imageimage-text-to-text1K<n<10K22 likes2.3k downloads1y agoHugging Face19keypa /vision-adapter-embeddings Vision Adapter MoonViT Embeddings Precomputed visual embeddings used to train lightweight vision→LLM projectors without re-running a vision tower: each row is the frozen MoonViT-V2 output for one training image, stored as raw bfloat16 bytes. Shards: 103 Parquet files (data/emb_0000.parquet … data/emb_0102.parquet), 1360 rows each, ~139k rows total, ~0.9 TB. Schema per row: column type meaning key string embedding id, embeddings/<sha1[:20]>.pt; matches emb in… See the full description on the dataset page: https://huggingface.co/datasets/keypa/vision-adapter-embeddings.textimage-to-text100K<n<1M0 likes2k downloads5d agoHugging Face20Vision-Flan /vision-flan_191-task_1k 🚀 Vision-Flan Dataset vision-flan_191-task-1k is a human-labeled visual instruction tuning dataset consisting of 191 diverse tasks and 1,000 examples for each task. It is constructed for visual instruction tuning and for building large-scale vision-language models. Paper or blog for more information: https://github.com/VT-NLP/MultiInstruct/ https://vision-flan.github.io/ Paper coming soon 😊 Citation Paper coming soon 😊. If you use Vision-Flan, please use the… See the full description on the dataset page: https://huggingface.co/datasets/Vision-Flan/vision-flan_191-task_1k.imagevisual-question-answering100K<n<1M22 likes1.9k downloads3y agoHugging Face21nyu-visionx /pisa-experiments Pisa Experiments This repository contains the PisaBench, training data, model checkpoints, introduced in PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop. PisaBench Real World Videos We curate a dataset comprising 361 videos demonstrating the dropping task.Each video begins with an object suspended by an invisible wire in the first frame. We cut the video clips to begin as soon as the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/pisa-experiments.n<1K2 likes1.9k downloads2y agoHugging Face22hf-vision /chest-xray-pneumoniaDataset Summary The dataset is organized into 3 folders (train, test, val) and contains subfolders for each image category (Pneumonia/Normal). There are 5,863 X-Ray images (JPEG) and 2 categories (Pneumonia/Normal). Chest X-ray images (anterior-posterior) were selected from retrospective cohorts of pediatric patients of one to five years old from Guangzhou Women and Children’s Medical Center, Guangzhou. All chest X-ray imaging was performed as part of patients’ routine clinical care. For… See the full description on the dataset page: https://huggingface.co/datasets/hf-vision/chest-xray-pneumonia.image1K<n<10K10 likes1.7k downloads3y agoHugging Face23Harvard-Edge /Wake-Vision-Train-Largeimage1M<n<10M2 likes1.7k downloads2y agoHugging Face24sasha /prof_images_blip__SG161222-Realistic_Vision_V1.4 Dataset Card for "prof_images_blip__SG161222-Realistic_Vision_V1.4" More Information needed image10K<n<100K0 likes1.6k downloads3y agoHugging Face25nyu-visionx /vstat VSTAT: Visual State Tracking Benchmark VSTAT is a video-based benchmark for evaluating the visual state tracking capability of Multimodal Large Language Models (MLLMs). It contains 834 video clips paired with 1,500 questions whose answers cannot be inferred from any single keyframe or short segment. Dataset Composition Split Videos Questions synthetic 450 550 self_recorded 80 100 youtube 304 850 Total 834 1,500 Files… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/vstat.tabularn<1K6 likes1.6k downloads2mo agoHugging Face26xDAN-Vision /Cambrian10M_For_Mantistext1M<n<10M1 likes1.5k downloads2y agoHugging Face27Harvard-Edge /Wake-Vision Dataset Card for Wake Vision Dataset Description "Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally, it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age, subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.imageimage-classification1M<n<10M11 likes1.4k downloads11mo agoHugging Face28Senqiao /VisionThink-Smart-Train VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning Senqiao/VisionThink-Smart-Train This is the training dataset used for our Efficient Reasoning VLM on general VQA tasks.VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning [Paper] Senqiao Yang, Junyi Li, Xin Lai, Bei Yu, Hengshuang Zhao, Jiaya Jia Highlights Our VisionThink leverages reinforcement learning to autonomously learn whether… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/VisionThink-Smart-Train.image10K<n<100K1 likes1.4k downloads1y agoHugging Face29xDAN-Vision /Docmatix_For_Mantistext1M<n<10M0 likes1.3k downloads2y agoHugging Face30xDAN-Vision /Websight_Mantis_Datatext1M<n<10M1 likes1.3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.