datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval.
multimodal_textbook
Multimodal-Textbook-6.5M
Overview
This dataset is for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining", containing 6.5M images interleaving with 0.8B text from instructional videos.
It contains pre-training corpus using interleaved image-text format. Specifically, our multimodal-textbook includes 6.5M keyframesextracted from instructional videos, interleaving with 0.8B ASR texts.
All the images and text are extracted from online… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/multimodal_textbook.propagator-multimodal-pretraining-data
Propagator Multimodal Pretraining Data
This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format.
This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout.
Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.multimodal_meme_classification_singapore
Dataset Card for Offensive Memes in Singapore Context
Dataset Details
Dataset Description
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards.
Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.multimodal-privacy
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
Recent advances in multi-modal Large Language Models (M-LLMs) have demonstrated a powerful ability to synthesize implicit information from disparate sources, including images and text. These resourceful data from social media also introduce a significant and underexplored privacy risk: the inference of sensitive personal attributes from seemingly daily media content. However, the lack of benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/xaddh/multimodal-privacy.McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval.
IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.STXBP1_PubMed_Central_Multimodal_Dataset
STXBP1 PubMed Central Multimodal Dataset v2 (12-13-2025)
A comprehensive multimodal dataset for training vision-language models on biomedical scientific literature, with focus on STXBP1-related neurological research.
🆕 Version 2 Updates (December 2025)
497,360 training examples (up from ~31K)
170,591 matched figure-image pairs (99.7% match rate)
Full captions preserved (no truncation)
Multiple training formats for different use cases
Validated response lengths for… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/STXBP1_PubMed_Central_Multimodal_Dataset.MultimodalMathBenchmarks
MultimodalMathBenchmarks
This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026).
It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs.
Canonical Upload Manifest
HF path
Local source
Count
Purpose
SharedMultimodalGrid.csv
SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.MultimodalUnlearningEvalBenchmark
🧠 Multimodal Unlearning Evaluation Benchmark
📌 Overview
This dataset provides evaluation outputs for studying metric inconsistency in multimodal machine unlearning.
It supports reproducibility of results in:
Metric Unreliability in Multimodal Machine Unlearning (NeurIPS 2026)
📊 Contents
File
Description
📄 multimodal_results.json
Results on VQA benchmarks (MLLMU-Bench, UnLOK-VQA, MMUBench)
📄 unimodal_results.json
CIFAR-10… See the full description on the dataset page: https://huggingface.co/datasets/neurips26/MultimodalUnlearningEvalBenchmark.panta_instruct_multi_modal_v1
Panta Instruct Multi-Modal v1
Dataset d'instructions multimodal en français : chaque exemple associe une question
(texte + parole + pictogrammes) à une réponse (texte + pictogrammes).
Colonnes
Colonne
Type
Description
audio
Audio (24 kHz, mono)
Enregistrement de la question (text_input)
text_input
string
Question / instruction
text_output
string
Réponse
pictos_input
list[string]
Identifiants des pictogrammes de la question
pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.long-context-multimodal-conversations-5k
long-context-multimodal-conversations-5k Synthetic Conversation & Compaction Dataset
Dataset Description
long-context-multimodal-conversations-5k is a fully synthetic dataset for training models to compact long multi-turn chat histories into structured, LLM-facing summaries.
The dataset has two primary artifacts:
Artifact
Description
Typical use
Conversations (conversations.jsonl)
Raw synthetic chat logs
Long-context modeling, compaction input… See the full description on the dataset page: https://huggingface.co/datasets/bravesoftware/long-context-multimodal-conversations-5k.pcb-defect-multi-modal-dataset
印制电路板焊接及装联缺陷多模态数据集
面向 PCB 焊接与装联缺陷识别、资料辅助根因分析和维修建议生成的中文多模态研究数据集。数据集将光学、扫描声学显微(SAM)和 X 射线等缺陷图像与来源标签、中文任务文本,以及可用的检测测量和试验工况关联,支持视觉语言模型的数据准备、指令微调实验和检索增强问答。
版本 v3 包含 5,000 条图像/缺陷区域样本、15,000 条任务指令,覆盖焊点连接质量、焊料分布与形态、焊接层结构及元件装联位置等方面的 14 类目标缺陷:
目标缺陷
含义与标注范围
虚焊
焊接界面未形成可靠的冶金结合,可能表现为接触不稳或间歇导通。
冷焊
热过程不足导致焊料未充分熔融或未形成合格连接,结合来源的热过程与外观信息标注。
少锡
焊料量或有效焊脚不足,焊接部位的填充或覆盖不充分。
多锡
焊料过量,形成过大的焊脚或明显堆积。
连锡
焊料连接本应相互隔离的端子或焊盘,形成焊料桥接。
开焊
应有焊接连接的部位缺少完整连接,或连接部位发生分离。
润湿不良… See the full description on the dataset page: https://huggingface.co/datasets/dezoe/pcb-defect-multi-modal-dataset.WebWorld-SFT-10k
WebWorld-SFT-10k
The Browser as a World Model for Self-Improving Web Code — a 10,000-example SFT release from the WebWorld training set.
Each example is a single-turn improvement of an HTML artifact that was accepted by the browser-issued certificate: the browser re-executed the candidate, target progress held, and every previously verified capability was preserved. Rejected trajectories are not in the export.
Why WebWorld
VLM-driven self-improvement of web code… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/WebWorld-SFT-10k.solarhive-community-solar-multimodal
SolarHive Community Solar Dataset
Canonical training corpus for the SolarHive family of fine-tuned Gemma 4 models. 1,727 rows (1,713 text + 14 image-grounded).
A combined text + sky-image training corpus for community solar energy intelligence. Built to fine-tune Gemma 4 into an AI energy advisor for residential solar microgrids — answering questions about production, storage, grid mix, weather impact, maintenance scheduling, and cross-source planning, with native… See the full description on the dataset page: https://huggingface.co/datasets/Truthseeker87/solarhive-community-solar-multimodal.multimodal-python-copilot-training-overview
Multimodal Datasets for Training Python Copilots from Source Code Analysis
Welcome to the matlok multimodal python copilot training datasets. This is an overview for our training and fine-tuning datasets found below:
~2.3M unique source coding rows
1.1M+ instruct alpaca yaml text rows updated bi-weekly
~923K png knowledge graph images with alpaca text description
~334K mp3s over ~2 years of continuous audio playtime
requires 1.5 TB storage on disk
Please reach out if you find an… See the full description on the dataset page: https://huggingface.co/datasets/matlok/multimodal-python-copilot-training-overview.fashion-stylist-multimodal
👗 Fashion Stylist Multimodal Dataset
A synthetic multimodal dataset pairing structured fashion metadata, styled outfit text, and AI-generated portraits.
🎯 Overview
This dataset contains ~1,000 synthetic fashion-styling profiles, each combining:
🧬 Structured demographic & style metadata
📝 A styled outfit description with e-commerce search queries
🖼️ A generated 512×512 studio-style portrait of a fictional person wearing the outfit
The dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/lihicarmeli/fashion-stylist-multimodal.multimodal_rag_complex_table_extractor_teaser
🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.gemma4-multimodal-recipe-dataset
🍳 Gemma 4 Multimodal Recipe & Food Dataset
A balanced, high-density multimodal dataset curated specifically for fine-tuning compact vision-language models (such as gemma-4-e2b-it) for visual food recognition, recipe generation, and dietary recommendation.
🔗 Upstream & Source Datasets
This dataset was created by cleaning, reformatting, and synthesizing samples across the following 5 Hugging Face sources:
Dataset
Modality
Role in Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/alst10/gemma4-multimodal-recipe-dataset.svg-multimodal-rubrics
SVG Multimodal Rubrics
A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects.
Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions.
Overview
Item
Details
Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.MultiModalDataset
Dataset Card for MultiModal Dataset
Dataset Description
Dataset Summary
MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation.
The dataset is organized into three subsets:
fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.multimodal_rewardbench
Dataset Card for Multimodal RewardBench
🏆 Dataset Attribution
This dataset is created by Yasunaga et al. (2025).
📄 Paper: Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
💻 GitHub Repository: https://github.com/facebookresearch/multimodal_rewardbench
I have downloaded the dataset from the GitHub repo and only modified the "Image" attribute by converting file paths to datasets.Image() for easier integration with 🤗… See the full description on the dataset page: https://huggingface.co/datasets/syhuggingface/multimodal_rewardbench.multimodal-vision-ocr-document-parsing-2026
📐 Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026)
This repository provides the official 100-sample production teaser of the Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights Vision-Language Models (Qwen2-VL, Pixtral-12B, Llama-3.2-Vision, ColPali) on dense document parsing, normalized spatial bounding boxes (<box>[ymin, xmin, ymax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-ocr-document-parsing-2026.ai-code-multimodal-fr
Dataset : IA Generation de Code, IA Multimodale & Small Language Models (FR)
Description
Dataset francophone couvrant les outils d'assistance au codage par IA, l'IA multimodale, les Small Language Models (SLM) et GraphRAG.
Ce dataset est concu pour la recherche, la formation et le developpement d'applications dans le domaine de l'IA appliquee au developpement logiciel et a la cybersecurite.
Articles couverts
IA pour la Generation de Code : Copilot, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-fr.lumos_multimodal_ground_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_ground_iterative.ai-code-multimodal-en
Dataset: AI Code Generation, Multimodal AI & Small Language Models (EN)
Description
English dataset covering AI coding assistants, multimodal AI, Small Language Models (SLMs), and GraphRAG.
This dataset is designed for research, training, and application development in the field of AI applied to software development and cybersecurity.
Articles Covered
AI Code Generation: Copilot, Cursor, Claude Code - Comparison of 12 leading AI coding assistants
Computer… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-en.lumos_multimodal_plan_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_plan_iterative.
