datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OraRL-Data
OraRL-Data
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code]
We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL.
It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.cktformer-dataset
CircuitFormer Dataset
This dataset contains a collection of 33,889 analog circuit netlists, images, and metadata collected from 62 textbooks.
Directory Structure
The dataset is organized into the following flattened directories (no train/test subfolders):
images/: Circuit diagrams (PNG format).
metadata/: JSON files containing circuit descriptions, names, and source information.
netlists/: SPICE netlist files (.cir).
Index Files
The dataset uses JSONL (JSON… See the full description on the dataset page: https://huggingface.co/datasets/touhid314/cktformer-dataset.seethrough3d-data
SeeThrough3D Dataset
Project Page | Paper | GitHub
This is the training dataset for the CVPR 2026 🎉 paper SeeThrough3D: Occlusion Aware 3D-Control in Text-to-Image Generation.
SeeThrough3D is a model for 3D layout-conditioned generation that explicitly models occlusions. This dataset consists of diverse multi-object scenes with strong inter-object occlusions, using an occlusion-aware 3D scene representation (OSCR) where objects are depicted as translucent 3D boxes.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/va1bhavagrawa1/seethrough3d-data.robomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.ManipDreamer3D_data
Dataset of paper ManipDreamer3D
This repository contains the dataset for the paper ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory.
Data Structure
The dataset is organized as follows:
manipdreamer3d_data/
├── 000000/ # data processed from bridge-v1
├── 000000_v2/ # data processed from bridge-v2
│ ├── data.json # contains gripper state, camera params, etc.
│ ├── depth_0000.png… See the full description on the dataset page: https://huggingface.co/datasets/myendless/ManipDreamer3D_data.CADBench-Extended-Multimodal-Dataset
Dataset Card
Dataset Description
CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics.
Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.Drivegpt4_raw_datatransientangelo_datasetInternVL-Chat-V1-2-SFT-Data
Data Card for InternVL-Chat-V1-2-SFT-Data
Overview
Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.ship-dataset
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning
ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels).
Quick reference
Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.qev-data
QEV data
Historical training snapshots for QEV, the LAYA-inspired Qwen3.5-2B decision model.
Model.
Configurations overlap. Do not concatenate them or assume independent test sets.
ZIPs in corpora/ use qev-<stage>.zip and contain eligible
original records, images and license notices. Extracted stage folders and original
record/source IDs preserve the recorded training provenance.
Viewer rows expose request/target schemas as JSON strings; parse with json.loads.
IDs, group IDs… See the full description on the dataset page: https://huggingface.co/datasets/ken-jo/qev-data.OPIS-dataset
OPIS Dataset
OPIS is an input-grounded benchmark for evaluating multi-object memory in
image-to-video and camera-conditioned video world models. It anchors every
evaluation to the fixed object instances visible in the initial image, rather
than to a generated history or a prescribed reference video.
The benchmark is designed to measure whether a generated rollout preserves the
objects established by its input. Its evaluator reports three complementary
dimensions: Presence… See the full description on the dataset page: https://huggingface.co/datasets/Kirito-Lab/OPIS-dataset.gemma4-serving-bench-data
Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data
Test data, charts, and the running research log from an autonomous research
loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via
llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop
summarizes findings, proposes a goal, tests it end-to-end, documents success or
failure, and publishes here + to GitHub.
Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K
ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.Reason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.FloodNet_2021-Track_2_Dataset_HF
FloodNet: High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding
This is the HF-hosted version of FloodNet.
The FloodNet 2021: A High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding provides high-resolution UAS imageries with detailed semantic annotation regarding the damages. To advance the damage assessment process for post-disaster scenarios, the authors of the dataset presented a unique challenge considering classification, semantic… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/FloodNet_2021-Track_2_Dataset_HF.medical-prescription-dataseteasyr1-agent-grounding-datamedical-prescription-datasetCabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-Behavior-Dataset.tdtu_vqa_dataset_herb
TDTU VQA Dataset — Vietnamese Medicinal Herbs 🌿
Dataset Description
TDTU VQA Dataset Herb is a Vietnamese Visual Question Answering (VQA) dataset focused on medicinal plants and herbs. It was developed for scientific research at Ton Duc Thang University (TDTU), with the goal of advancing AI models capable of recognizing and answering questions about Vietnamese medicinal herbs.
Homepage: Hugging Face Dataset
Repository: azan100an/tdtu_vqa_dataset_herb
Point of Contact:… See the full description on the dataset page: https://huggingface.co/datasets/azan100an/tdtu_vqa_dataset_herb.Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.nornikel-metallurgy-vl-dataset
Nornikel Metallurgy VL Dataset (SFT / DPO / GRPO)
Датасет для дообучения мультимодальной модели Qwen3-VL по схеме
SFT → DPO → GRPO в предметной области металлургии, горного дела и
обогащения полезных ископаемых. Построен из корпуса технических документов
(PDF-книги/сборники, DOCX-отчёты, PPTX-презентации, XLSX-таблицы) и
изображений (схемы, диаграммы, таблицы).
Конфигурации (config_name)
config
train
validation
назначение
sft
111 351
12 372… See the full description on the dataset page: https://huggingface.co/datasets/brics-edtech/nornikel-metallurgy-vl-dataset.pamela
PAM∃LA
Personalizing Text-to-Image Generation to Individual Taste
Anonymous submission — author and affiliation details withheld during review.
PAM∃LA is a dataset of AI-generated images rated by human participants for aesthetic quality, built specifically for personalization research. It pairs each rating with rich participant demographics and image metadata, enabling research on personalized aesthetic prediction, demographic variation in visual preference, and reward modelling for… See the full description on the dataset page: https://huggingface.co/datasets/pamela-dataset/pamela.pcqa_dataset_projectionsVBVR-Bench-Data
VBVR: A Very Big Video Reasoning Suite
Overview
Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture,
enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality.
Systematically studying video reasoning and its scaling behavior suffers from a lack of… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Bench-Data.safer-ar-dataPAD3-Dataset-Revisi-Fixed-TRL
PAD3-Dataset-Revisi-Fixed (TRL chat format)
Conversational (TRL / SFT) dataset for age-rating classification of images.
Structure
.
├── metadata.jsonl # one TRL chat record per line
└── images/
├── Semua_Umur/000000.jpg
├── 7_/000000.jpg
├── 13_/000000.jpg
├── 15_/000000.jpg
├── 18_/000000.jpg
└── Konten_Terlarang/000000.jpg
Images are split into per-rating subfolders to stay under the 10,000-files-per-folder limit.… See the full description on the dataset page: https://huggingface.co/datasets/capstone-pad3/PAD3-Dataset-Revisi-Fixed-TRL.aitf-dfk3-vlm-dataset-jsonlvila-q-data-train
