datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.MM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.MLLMU-Bench
Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
Abstract
Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/MLLMMU/MLLMU-Bench.UCITUnofficial training-ready fork of HaiyangGuo/UCIT
MLLM-as-a-JudgeEarthScience-MLLM-20K
EarthScience-MLLM-20K
A unified JSONL package for multimodal large-model training across three Earth-science domains:
Meteorology from ZhanxiangHua/WeatherQA_SFT.
Geography / map QA from HuggingFaceM4/the_cauldron config mapqa.
Remote-sensing common-sense QA + grounding/detection from xiang709/VRSBench.
The package intentionally excludes segmentation-style targets. Each JSONL line is one training/evaluation unit.
Files
train.jsonl: 20000 examples.
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/moTcream/EarthScience-MLLM-20K.Domain40ksarab
Sarab Dataset
The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination
evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD
benchmark. Code and evaluation scripts are on
GitHub.
One real item from each mode, with the question in Arabic and English, the expected answer, and an actual model reply (a tick marks a correct first word, a cross a wrong one).
What this is
A human-captioned pool of Arabic Cultural Visual… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.OmniParsingBench
🤗 Model | 📑 Technical Report | 💻 GitHub
OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities.
Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.MLLMGuard
MLLMGuard
MLLMGuard is a multi-dimensional safety evaluation suite for MLLMs, including a bilingual
image-text evaluation dataset, inference utilities, and a set of lightweight evaluators.
Quick Links
arXiv Paper
Github Repository
Acquisition of Datasets
The datasets corresponding to the results in the paper are unmasked versions. You can obtain the datasets by filtering out the form. The review results will be sent to your email within 1-2 business days.
VTCBench
Dataset Card for VTCBench
Vision-Text Compression Benchmark (VTCBench)
revisits Needle-In-A-Haystack (NIAH)
from a VLM's perspective by converting long context into rendered images.
This benchmark tests VLM's ability to OCR, retrieve, aggregate, infer, and
memorize long context as images. Specifically, this benchmark includes 3 tasks:
Retrieval: Vision-NIAH VQA task for information retrieval and aggregation.… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/VTCBench.mllm_attack_dataMLLMKC-datasetsDCLMLLM_agriculture_evaluation
Evaluation methodology dataset for multimodal LLM-based system with geospatial augmented context for agricultural plot-level assistance
This dataset was built for evaluating context-amplified multimodal Large Language Models (LLM) within the agricultural domain. It is supplementary material for a manuscript currently under peer review.
Data Structure
The dataset contains the following columns visible in the viewer table above:
QueryID (string): The unique… See the full description on the dataset page: https://huggingface.co/datasets/Juancc-ctic/MLLM_agriculture_evaluation.Comic-9K
Comic-9K
Image
Extracting all images.
cat images.tar.gz.aa images.tar.gz.ab images.tar.gz.ac images.tar.gz.ad images.tar.gz.ae > images.tar.gz
tar xvzf images.tar.gz
Summary
We provide human-written plot synopsis.
summary.jsonl
GeneralScience-MLLM-22K
GeneralScience-MLLM-22K
Dataset Summary
GeneralScience-MLLM-22K is a unified general-science multiple-choice QA collection built from local snapshots of SciQ, AI2 ARC, and ScienceQA. It follows the same release style as a subject-specific MLLM dataset: every sample is stored as one JSONL record, text-only and image-text examples share one schema, and ScienceQA images are exported as standalone files referenced by relative paths.
The release contains 22,661… See the full description on the dataset page: https://huggingface.co/datasets/gineven/GeneralScience-MLLM-22K.Spatial-MLLM-DataThis repository contains the some datasets used for Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.
spatial-mllm-datamllm_cl_vizwizmllm_cl_textvqaMLLM_testMLLM-Fabric
🧵 MLLM-Fabric: Multimodal LLM-Driven Robotic Framework for Fabric Sorting and Selection
📄 Overview
This is the official repository for the paper:
MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection
Accepted to IEEE Robotics and Automation Letters (RA-L)
🏫 About This Work
This work is from the Robot-Assisted Living LAboratory (RALLA) at the University of York, UK.
🧵 Fabric Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/EuniceF/MLLM-Fabric.spatial-mllm-qwen72b-kitMLLM3-textcaps-scienceqa-vqav2S-MLLMUn-data
S-MLLMUn
Official implementation of the ECCV 2026 paper
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
Paper
·
Code
·
Dataset
·
LLaVA-OneVision Original Model
·
Qwen2.5-VL Original Model
S-MLLMUn Bench
This repository contains the parquet files used by S-MLLMUn Bench for selective multimodal large language model unlearning.
Contents
ft_data: fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/ZhenZeng/S-MLLMUn-data.MLLMGuard
MLLMGuard
MLLMGuard is a multi-dimensional safety evaluation suite for MLLMs, including a bilingual
image-text evaluation dataset, inference utilities, and a set of lightweight evaluators.
Quick Links
arXiv Paper
Github Repository
Acquisition of Datasets
The datasets corresponding to the results in the paper are unmasked versions. You can obtain the datasets by filtering out the form. The review results will be sent to your email within 1-2 business… See the full description on the dataset page: https://huggingface.co/datasets/Shaozhiyuan/MLLMGuard.localize-indoor
Elliot Localize Indoor / 3D and depth
Upstream training splits; known explicitly identified test/eval rows excluded. Cross-dataset benchmark overlap is not guaranteed. Published as a raw, manually gated release; source annotation caveats remain.
Task views reuse original image archives. 2D coordinates are normalized 0–1000. Native 3D sidecars preserve original camera-space XYZ and camera calibration; they are not normalized to 0–1000. Within each query targets are sorted… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-indoor.drivelm_cleaned
drivelm_cleaned
The drivelm__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
2,804
QA turns
71,425
answers rewritten by the cleaning pass
18,520
QA created by the cleaning pass (new_qa)
13,889 (19.4%)
shards
8
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/drivelm_cleaned.localize-gui
Elliot Localize GUI
Upstream training splits; known explicitly identified test/eval rows excluded. Cross-dataset benchmark overlap is not guaranteed. Published as a raw, manually gated release; source annotation caveats remain.
Task views reuse original image archives. Coordinates are normalized 0–1000. Within each query targets are sorted left-to-right then top-to-bottom.
HF preview configs contain 10 examples per view, not the complete training split. Full training uses… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-gui.
