datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.FGVC_Aircraft_test
Dataset Card for "FGVC_Aircraft_test"
More Information needed
FGVC_Aircraft_train
Dataset Card for "FGVC_Aircraft_train"
More Information needed
adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.walton-multimodal-cold-start-r1-format
walton-multimodal-cold-start-r1-format
WaltonFuture/Multimodal-Cold-Start converted to multimodal-open-r1-8k-verified format with filtering
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/walton-multimodal-cold-start-r1-format.rand-1m-multimodalmultimodal-ai-taxonomy
Multimodal AI Taxonomy
A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities.
Dataset Description
This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to:
Navigate the complex landscape of multimodal AI models
Filter models by specific input/output modality combinations
Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.KomdigiITS-DFK3-Multimodalai-code-multimodal-fr
Dataset : IA Generation de Code, IA Multimodale & Small Language Models (FR)
Description
Dataset francophone couvrant les outils d'assistance au codage par IA, l'IA multimodale, les Small Language Models (SLM) et GraphRAG.
Ce dataset est concu pour la recherche, la formation et le developpement d'applications dans le domaine de l'IA appliquee au developpement logiciel et a la cybersecurite.
Articles couverts
IA pour la Generation de Code : Copilot, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-fr.lumos_multimodal_ground_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_ground_iterative.Viet-multimodal-open-r1-8k-verifiedai-code-multimodal-en
Dataset: AI Code Generation, Multimodal AI & Small Language Models (EN)
Description
English dataset covering AI coding assistants, multimodal AI, Small Language Models (SLMs), and GraphRAG.
This dataset is designed for research, training, and application development in the field of AI applied to software development and cybersecurity.
Articles Covered
AI Code Generation: Copilot, Cursor, Claude Code - Comparison of 12 leading AI coding assistants
Computer… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-en.FGVC_Aircraft_test_facebook_opt_350m_Attributes_Caption_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_350m_Attributes_Caption_ns_3333"
More Information needed
FGVC_Aircraft_test_facebook_opt_350m_Visclues_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_350m_Visclues_ns_3333"
More Information needed
lumos_multimodal_plan_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_plan_iterative.FGVC_Aircraft_test_facebook_opt_2.7b_Visclues_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_2.7b_Visclues_ns_3333"
More Information needed
AI4Math_MathVista-multimodal-rollout8FGVC_Aircraft_test_facebook_opt_1.3b_Visclues_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_1.3b_Visclues_ns_3333"
More Information needed
FGVC_Aircraft_test_facebook_opt_350m_Attributes_Caption_ns_3333_random
Dataset Card for "FGVC_Aircraft_test_facebook_opt_350m_Attributes_Caption_ns_3333_random"
More Information needed
AI4Math_MathVista-multimodal_7B-rollout8africa-synth-aid-flows-radar-multimodal-edge-all
Multi-Modal Edge Sensor Signatures (Radar + IMU + Acoustic) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-radar-multimodal-edge-all.FGVC_Aircraft_test_facebook_opt_1.3b_Attributes_Caption_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_1.3b_Attributes_Caption_ns_3333"
More Information needed
FGVC_Aircraft_test_facebook_opt_2.7b_Attributes_Caption_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_2.7b_Attributes_Caption_ns_3333"
More Information needed
multimodal-open-r1-8192-filtered-mid-ic
multimodal-open-r1-8192-filtered-mid-ic
Original dataset structure preserved, filtered by token length and image quality
Dataset Description
This dataset was processed using the data-preproc package for vision-language model training.
Processing Configuration
Base Model: Qwen/Qwen2.5-7B-Instruct
Tokenizer: Qwen/Qwen2.5-7B-Instruct
Sequence Length: 16384
Processing Type: Vision Language (VL)
Dataset Features
input_ids: Tokenized input sequences… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/multimodal-open-r1-8192-filtered-mid-ic.FGVC_Aircraft_test_facebook_opt_2.7b_Visclues_ns_3333_random
Dataset Card for "FGVC_Aircraft_test_facebook_opt_2.7b_Visclues_ns_3333_random"
More Information needed
FGVC_Aircraft_test_facebook_opt_6.7b_Attributes_Caption_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_6.7b_Attributes_Caption_ns_3333"
More Information needed
FGVC_Aircraft_test_facebook_opt_6.7b_Visclues_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_6.7b_Visclues_ns_3333"
More Information needed
