datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.AgMMU_v1
AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
Aruna Gauba1,2,5*,
Irene Pi1,3,5*,
Yunze Man1,4,5†,
Ziqi Pang1,4,5†,
Vikram S. Adve1,4,5,
Yu-Xiong Wang1,4,5
1University of Illinois at Urbana-Champaign, 2Rice University, 3Carnegie Mellon University
4AIFARMS, 5Center for Digital Agriculture at UIUC
Introduction
AgMMU is a challenging real‑world benchmark for evaluating and advancing… See the full description on the dataset page: https://huggingface.co/datasets/AgMMU/AgMMU_v1.llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates.
The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero
LLaVA-1.5-665K-Instructions
This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences.
The images are in train_split/*.tars and the text sequences are in jsons:
llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.VISTA-400K
VISTA-400K
This repo contains all subsets for VISTA-400K. VISTA is a video spatiotemporal augmentation method that generates long-duration and high-resolution video instruction-following data to enhance the video understanding capabilities of video LMMs.
This repo is under construction. Please stay tuned.
🌐 Homepage | 📖 arXiv | 💻 GitHub | 🤗 VISTA-400K | 🤗 Models | 🤗 HRVideoBench
Video Instruction Data Synthesis Pipeline
VISTA leverages insights from… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VISTA-400K.MECAT-QAMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-Caption (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-QA.train_raw_video
ShareGPTVideo Raw ActivityNet Videos for Train data
All dataset and models can be found at ShareGPTVideo.
Contents:
Due to our scene split, we provide our processed activityNet videos corresponding to test frames in
train video frames
the processing script is process_activitynet.py
StreamGaze_v2
StreamGaze Dataset
StreamGaze is a comprehensive streaming video benchmark for evaluating MLLMs on gaze-based QA tasks across past, present, and future contexts.
Companion dataset: The EgoGazeVQA dataset is hosted separately at Peanuttoad/gaze_dataset.
📁 Dataset Structure
streamgaze/
├── metadata/
│ ├── egtea.csv # EGTEA fixation metadata
│ ├── egoexolearn.csv # EgoExoLearn fixation metadata
│ └── holoassist.csv # HoloAssist… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/StreamGaze_v2.MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.azerbaijan-court-data
Azerbaijan Court System Dataset
The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations.
Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale.
Quick Start
Load with Hugging Face datasets
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.ocr-vqa-200k_imagesImage collections for OCR-VQA-200K.
Image size: 208,467.
VGAVS-Datasetbase_paper:
arxiv
test_raw_video_data
ShareGPTVideo Raw Videos for Testing data
All dataset and models can be found at ShareGPTVideo.
Contents:
In case of need, this contains raw videos corresponding to test frames in
Test video frames
EgoExoBench_MCQ
🧠 EgoExoBench: A Cross-Perspective Video Understanding Benchmark
Dataset Summary
EgoExoBench is a benchmark designed to evaluate cross-perspective understanding capabilities of multimodal large models (MLLMs).It contains synchronized and asynchronous egocentric (first-person) and exocentric (third-person) video pairs, along with multiple-choice questions that assess semantic alignment, viewpoint association, and temporal reasoning between the two perspectives.… See the full description on the dataset page: https://huggingface.co/datasets/Heleun/EgoExoBench_MCQ.WearVQA
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world Scenarios
Paper: WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenariosAuthors: Eun Chang*, Zhuangqun Huang*, Yiwei Liao*, Sagar Ravi Bhavsar*, Amogh Param, Tammy Stark, Adel Ahmadyan, Xiao Yang, Jiaqi Wang, Ahsan Abdullah, Giang Nguyen, Akil Iyer, David Hall, Elissa Li, Shane Moon, Nicolas Scheffer, Kirmani Ahmed, Babak Damavandi, Rakesh… See the full description on the dataset page: https://huggingface.co/datasets/tonyliao-meta/WearVQA.madqa-training
Chrisyichuan/madqa-training
MADQA document QA contrastive training data with hard negatives.
Contents
madqa_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 1840
unique images: 3598
avg… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/madqa-training.peek_vqa
Peek VQA Dataset Card
Dataset details
This dataset contains 2M image-QA pairs used to fine-tune PEEK VLM, a VLM for robotics that answers "where policies should focus" and "what policies should do".
Given an instruction (<quest></quest>), the task is to predict a unified point-based representation corresponding to 1) a path guiding the robot end-effector in what actions to take (TRAJECTORY), and 2) a set of task-relevant masking points that show where to focus on (MASK).… See the full description on the dataset page: https://huggingface.co/datasets/memmelma/peek_vqa.moca-colpali-training
Chrisyichuan/moca-colpali-training
MOCA ColPali contrastive training data with hard negatives.
Contents
moca_colpali_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 118195
unique… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-colpali-training.moca-visrag-syn-training
Chrisyichuan/moca-visrag-syn-training
MOCA VisRAG synthetic-split contrastive training data with hard negatives.
Contents
moca_visrag_syn_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 239206… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-syn-training.moca-visrag-ind-training
Chrisyichuan/moca-visrag-ind-training
MOCA VisRAG independent-split contrastive training data with hard negatives.
Contents
moca_visrag_ind_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 122752… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-ind-training.OKReddit-Visionary
Dataset Summary
OKReddit Visionary is a collection of 50 GiB (~74K pairs) of image Question & Answers. This dataset has been prepared for research or archival purposes.
Curated by: KaraKaraWitch
Funded by: Recursal.ai
Shared by: KaraKaraWitch
Special Thanks: harrison (Suggestion)
Language(s) (NLP): Mainly English.
License: Refer to Licensing Information for data license.
Dataset Sources
Source Data: Academic Torrents by (stuck_in_the_matrix, Watchful1… See the full description on the dataset page: https://huggingface.co/datasets/recursal/OKReddit-Visionary.
