Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes8.7k downloads3mo agoHugging Face02Multilingual-Multimodal-NLP /IfEvalCode-testsettextn<1K2 likes4.3k downloads1y agoHugging Face03Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes3.8k downloads2y agoHugging Face04electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.6k downloads2mo agoHugging Face05LianeMarilin /CADBench-Extended-Multimodal-Dataset Dataset Card Dataset Description CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics. Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.3dimage-to-textn<1K2 likes2.1k downloads1mo agoHugging Face06Multilingual-Multimodal-NLP /IfEvalCode-Instructtext1K<n<10K2 likes962 downloads1y agoHugging Face07FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes626 downloads2y agoHugging Face08Multilingual-Multimodal-NLP /McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval. texttext-generation10K<n<100K39 likes452 downloads2y agoHugging Face09f13rnd /multimodal-example Multimodal Example Dataset Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift. Structure ├── train.jsonl # 10 training samples ├── test.jsonl # 2 validation samples ├── images/ # All referenced images (400x300 JPEG) │ ├── dog_portrait.jpg │ ├── forest_river.jpg │ ├── laptop_desk.jpg │ ├── mountain_lake.jpg │ ├── ocean_rocks.jpg │ ├── coffee_cup.jpg │ ├── bookshelf.jpg │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.imagen<1K0 likes385 downloads6mo agoHugging Face10Multilingual-Multimodal-NLP /MdEval MDEVAL: Massively Multilingual Code Debugging Official repository for our paper "MDEVAL: Massively Multilingual Code Debugging" 🏠 Home Page • 📊 Benchmark Data • 🏆 Leaderboard Introduction MDEVAL is a massively multilingual debugging benchmark covering 20 programming languages with 3.9K test samples and three tasks focused on bug fixing. It substantially pushes the limits of code LLMs in multilingual… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/MdEval.text10K<n<100K5 likes379 downloads9mo agoHugging Face11June30916 /multimodality-poc-llama31-ruler16k Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K Raw pre-RoPE query and hidden-state tensors captured during prefill, used to study whether the per-(layer, kv_head) query distribution is unimodal Gaussian (the assumption underpinning Expected Attention's MGF closed-form in kvpress). What's in here 65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5 prompts). Each file (~414 MB) contains: field dtype shape meaning hidden float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.tabularfeature-extractionn<1K0 likes227 downloads5mo agoHugging Face12huanngzh /DeepFashion-MultiModal-Parts2Whole DeepFashion MultiModal Parts2Whole Dataset Details Dataset Description This human image dataset comprising about 41,500 reference-target pairs. Each pair in this dataset includes multiple reference images, which encompass human pose images (e.g., OpenPose, Human Parsing, DensePose), various aspects of human appearance (e.g., hair, face, clothes, shoes) with their short textual labels, and a target image featuring the same individual (ID) in the same outfit… See the full description on the dataset page: https://huggingface.co/datasets/huanngzh/DeepFashion-MultiModal-Parts2Whole.imagetext-to-image10K<n<100K9 likes211 downloads2y agoHugging Face13hashmortar /multimodal-annual-reports Multimodal Annual Reports A document question-answering benchmark built from 20 complete corporate annual and integrated reports. It contains 595 English questions with reference answers, source evidence, original PDFs, reviewed HTML and Markdown representations, and figure/table crops. Questions require interpreting narratives, tables, and non-tabular visuals, including Japanese and French sources. MIT covers original benchmark contributions only. Source reports and their… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/multimodal-annual-reports.imagedocument-question-answering1K<n<10K0 likes195 downloads4d agoHugging Face14WeThink /WeThink-Multimodal-Reasoning-120K WeThink-Multimodal-Reasoning-120K Image Type Images data can be access from https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k Image Type Source Dataset Images General Images COCO 25,344 SAM-1B 18,091 Visual Genome 4,441 GQA 3,251 PISC 835 LLaVA 134 Text-Intensive Images TextVQA 25,483 ShareTextVQA 538 DocVQA 4,709 OCR-VQA5,142 ChartQA 21,781 Scientific & Technical GeoQA+ 4,813 ScienceQA 4,990 AI2D 1,812 CLEVR-Math 677… See the full description on the dataset page: https://huggingface.co/datasets/WeThink/WeThink-Multimodal-Reasoning-120K.textvisual-question-answering100K<n<1M7 likes189 downloads1y agoHugging Face15Multilingual-Multimodal-NLP /TableInstruct Citation @misc{wu2024tablebenchcomprehensivecomplexbenchmark, title={TableBench: A Comprehensive and Complex Benchmark for Table Question Answering}, author={Xianjie Wu and Jian Yang and Linzheng Chai and Ge Zhang and Jiaheng Liu and Xinrun Du and Di Liang and Daixin Shu and Xianfu Cheng and Tianzhen Sun and Guanglin Niu and Tongliang Li and Zhoujun Li}, year={2024}, eprint={2408.09174}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/TableInstruct.text10K<n<100K17 likes184 downloads2y agoHugging Face16superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face17fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes148 downloads22d agoHugging Face18Multilingual-Multimodal-NLP /AutoMemoryBench AutoMemoryBench State-Contract Evaluation for Auditable Agent Memory AutoMemoryBench evaluates whether an agent uses the right memory, and only the admissible memory, under a query-time state contract. Each executable contract partitions memory into required, admissible, and prohibited sets. Prohibited memories are typed as superseded, deleted, restricted, cross-namespace, or stale-tool. Relevance is not enough: remembered evidence must also be… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/AutoMemoryBench.textquestion-answering100K<n<1M0 likes124 downloads2mo agoHugging Face19multimodalart /agent-spaces-tracestabularn<1K0 likes95 downloads6mo agoHugging Face20RKB109 /multimodal-document-retrieval-20260911-dataset Multimodal Document Retrieval Baseline Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260911-dataset.textvisual-document-retrievaln<1K0 likes87 downloads29d agoHugging Face21yangjie-cv /WeThink_Multimodal_Reasoning_120K Dataset Card for WeThink Repository: https://github.com/yangjie-cv/WeThink Paper: https://arxiv.org/abs/2506.07905 Dataset Structure Question-Answer Pairs The WeThink_Multimodal_Reasoning_120K.jsonl file contains the question-answering data in the following format: { "problem": "QUESTION", "answer": "ANSWER", "category": "QUESTION TYPE", "abilities": "QUESTION REQUIRED ABILITIES", "refined_cot": "THINK PROCESS", "image_path": "IMAGE PATH"… See the full description on the dataset page: https://huggingface.co/datasets/yangjie-cv/WeThink_Multimodal_Reasoning_120K.text100K<n<1M12 likes86 downloads1y agoHugging Face22RKB109 /multimodal-document-retrieval-20260921-dataset Multimodal Document Retrieval Baseline Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260921-dataset.textvisual-document-retrievaln<1K0 likes86 downloads19d agoHugging Face23danielrosehill /multimodal-ai-taxonomy Multimodal AI Taxonomy A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities. Dataset Description This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to: Navigate the complex landscape of multimodal AI models Filter models by specific input/output modality combinations Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.textothern<1K0 likes65 downloads1y agoHugging Face24erobinson /repro-mm-deepresearch-a-simple-and-effective-multimodal-agentic-search-baseline-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes63 downloads2mo agoHugging Face25obaydata /svg-multimodal-rubrics SVG Multimodal Rubrics A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects. Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions. Overview Item Details Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.imagetext-generationn<1K0 likes55 downloads7mo agoHugging Face26LIAGM /DeepFashion-MultiModal-Parts2Whole DeepFashion MultiModal Parts2Whole Dataset Details Dataset Description This human image dataset comprising about 41,500 reference-target pairs. Each pair in this dataset includes multiple reference images, which encompass human pose images (e.g., OpenPose, Human Parsing, DensePose), various aspects of human appearance (e.g., hair, face, clothes, shoes) with their short textual labels, and a target image featuring the same individual (ID) in the same outfit… See the full description on the dataset page: https://huggingface.co/datasets/LIAGM/DeepFashion-MultiModal-Parts2Whole.imagetext-to-image10K<n<100K2 likes53 downloads2y agoHugging Face27aitf-komdigi /KomdigiITS-DFK3-Multimodalimage10K<n<100K0 likes53 downloads3mo agoHugging Face28RKB109 /multimodal-document-retrieval-20261001-dataset Multimodal Document Retrieval Baseline Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20261001-dataset.textvisual-document-retrievaln<1K0 likes51 downloads9d agoHugging Face29tongliuphysics /multimodalpragmatic Multimodal Pragmatic Jailbreak on Text-to-image Models Project page | Paper | Code The Multimodal Pragmatic Unsafe Prompts (MPUP) is a dataset designed to assess the multimodal pragmatic safety in Text-to-Image (T2I) models. It comprises two key sections: image_prompt, and text_prompt. Dataset Usage Downloading the Data To download the dataset, install Huggingface Datasets and then use the following command: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/tongliuphysics/multimodalpragmatic.tabulartext-to-image1K<n<10K3 likes46 downloads18d agoHugging Face30RKB109 /multimodal-document-retrieval-20260812-dataset Multimodal Document Retrieval Baseline Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260812-dataset.textvisual-document-retrievaln<1K0 likes44 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.