Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lin-Chen /MMStar MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub Dataset Details As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs' and LVLMs' training data. Therefore, we introduce MMStar: an elite vision-indispensible multi-modal benchmark, aiming to ensure each curated sample exhibits… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/MMStar.imagemultiple-choice1K<n<10K54 likes18k downloads3y agoHugging Face023dllm /MMScan-betatext1M<n<10M1 likes5k downloads2y agoHugging Face03ddwang2000 /MMSU [ICLR 2026] MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark Overview of MMSU MMSU (Massive Multi-task Spoken Language Understanding and Reasoning Benchmark) is a comprehensive benchmark for evaluating fine-grained spoken language understanding and reasoning in multimodal models. It systematically captures the variance of real-world linguistic phenomena in daily speech through 47 sub-tasks, including phonetics, prosody, rhetoric… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/MMSU.audioquestion-answering1K<n<10K16 likes3.2k downloads6mo agoHugging Face04coml /mmsulab DiscoPhon - Segmented MMS ulab v2 This dataset is a segmented version of espnet/mms_ulab_v2 using pyannote/segmentation-3.0. License and Acknowledgement Following espnet/mms_ulab_v2, this dataset is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 license. If you use this dataset, please cite the DiscoPhon paper @inproceedings{poli2026discophon, title = {{DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories… See the full description on the dataset page: https://huggingface.co/datasets/coml/mmsulab.audioaudio-to-audio1M<n<10M2 likes2.9k downloads1d agoHugging Face05RomeroLab-Duke /af3-mmseqs-db1 likes2.7k downloads6mo agoHugging Face06mmbench /MM-SpuBench MM-SpuBench Datacard Basic Information Title: The Multimodal Spurious Benchmark (MM-SpuBench) Description: MM-SpuBench is a comprehensive benchmark designed to evaluate the robustness of MLLMs to spurious biases. This benchmark systematically assesses how well these models distinguish between core and spurious features, providing a detailed framework for understanding and quantifying spurious biases. Data Structure: ├── data/images │ ├── 000000.jpg │ ├── 000001.jpg │… See the full description on the dataset page: https://huggingface.co/datasets/mmbench/MM-SpuBench.imagequestion-answering1K<n<10K2 likes2.4k downloads2y agoHugging Face07PKU-Alignment /MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use). Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.image1K<n<10K8 likes1.9k downloads2y agoHugging Face08EchoSafe-MLLM /MM-SafetyBench-plus-plus MM-SafetyBench++ Project Page | Paper | Code MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent. Dataset Summary For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.imageimage-text-to-text1K<n<10K2 likes1.8k downloads7mo agoHugging Face09RunsenXu /MMSI-Bench MMSI-Bench This repo contains evaluation code for the paper "MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence" 🌐 Homepage | 🤗 Dataset | 📑 Paper | 💻 Code | 📖 arXiv 🔔News 🔥[2025-10-23]: We added the normalized human response time for each MMSI-Bench sample and its difficulty level to our dataset on Hugging Face. 🔥[2025-06-18]: MMSI-Bench has been supported in the LMMs-Eval repository. ✨[2025-06-11]: MMSI-Bench was used for evaluation in the… See the full description on the dataset page: https://huggingface.co/datasets/RunsenXu/MMSI-Bench.imagequestion-answering1K<n<10K17 likes1.7k downloads1y agoHugging Face10Voxel51 /yuto-mms-multimodal Dataset Card for YUTO MMS Multimodal (MCAP) A FiftyOne build of YUTO MMS (York University Teledyne Optech Mobile Mapping System Dataset), a SLAM benchmark from the AUSM Lab at York University. This build repackages the source dataset's per-sequence raw sensor folders as time-synchronized MCAP recordings for FiftyOne's native multimodal dataset support (FiftyOne 1.19+). Each sample is one episode (one continuous drive), viewable in FiftyOne's tiled multimodal viewer with a… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/yuto-mms-multimodal.roboticsn<1K3 likes1.7k downloads2mo agoHugging Face11Cie1 /MMSearch-Plus MMSearch-Plus✨: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents Official repository for the paper "MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents". 🌟 For more details, please refer to the project page with examples: https://mmsearch-plus.github.io/. [🌐 Webpage] [📖 Paper] [🤗 Huggingface Dataset] [🏆 Leaderboard] 💥 News [2025.09.26] 🔥 We update the arXiv paperand release all MMSearch-Plus data samples in… See the full description on the dataset page: https://huggingface.co/datasets/Cie1/MMSearch-Plus.imagequestion-answeringn<1K2 likes1.3k downloads6mo agoHugging Face12zhaochenyang20 /mmsu-ci-2000audio1K<n<10K0 likes1.2k downloads6mo agoHugging Face13rbler /MMSI-Video-Bench MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence 🌐 Homepage | 📑 Paper | 📖 Code 🔔 News 🔥[2025-12]: Our MMSI-Video-Bench has been integrated into VLMEvalKit. 🔥[2025-12]: We released our paper, benchmark, and evaluation codes. 📊 Data Details All of our data is available on Hugging Face and includes the following components: 🎥 Video Data (videos.zip): Contains the video clip file (.mp4) corresponding to each sample. This… See the full description on the dataset page: https://huggingface.co/datasets/rbler/MMSI-Video-Bench.multiple-choice1K<n<10K6 likes1.1k downloads8mo agoHugging Face14rbler /MMScan-llava-form MMScan LLaVA-Form Data This repository provides the processed LLaVA-formatted dataset for the MMScan Question Answering Benchmark. Dataset Contents (1) All image data(Depth&RGB) is distributed in split ZIP archives. Please combine the split ZIP files into a single archive and extract the merged ZIP file using the following command: cat mmscan_val8.z* > mmscan_va.zip unzip mmscan_va.zip (2) Under ./annotations, we provide the MMScan Question Answering validation set with… See the full description on the dataset page: https://huggingface.co/datasets/rbler/MMScan-llava-form.video-text-to-text100K<n<1M1 likes1.1k downloads1y agoHugging Face15Yiwei-Ou /MMS-VPR MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments Overview MMS-VPR is the first large-scale multimodal street-level visual place recognition dataset featuring comprehensive integration of images, videos, and rich textual annotations with day–night coverage and a 7-year temporal span in dense pedestrian-only environments. MMS-VPR comprises 110,529 images and 2,527 video clips… See the full description on the dataset page: https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR.imageimage-classificationn<1K2 likes1.1k downloads5mo agoHugging Face16OmniSVG /MMSVG-IllustrationOmniSVG: A Unified Scalable Vector Graphics Generation Model ![Project Page] Dataset Card for MMSVG-Illustration Dataset Description This dataset contains SVG illustration examples for training and evaluating SVG models for text-to-SVG and image-to-SVG task. Dataset Structure Features The dataset contains the following fields: Field Name Description id Unique ID for each SVG svg SVG code (resized to 200×200, simplified with picosvg)… See the full description on the dataset page: https://huggingface.co/datasets/OmniSVG/MMSVG-Illustration.image100K<n<1M66 likes1k downloads10mo agoHugging Face17espnet /mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/espnet/mms_ulab_v2.audioaudio-to-audio10K<n<100K27 likes947 downloads2y agoHugging Face18OmniSVG /MMSVG-IcongatedOmniSVG: A Unified Scalable Vector Graphics Generation Model ![Project Page] Dataset Card for MMSVG-Icon Dataset Description This dataset contains SVG icon examples for training and evaluating SVG models for text-to-SVG and image-to-SVG task. Dataset Structure Features The dataset contains the following fields: Field Name Description id Unique ID for each SVG svg SVG code (resized to 200×200, simplified with picosvg) description… See the full description on the dataset page: https://huggingface.co/datasets/OmniSVG/MMSVG-Icon.image100K<n<1M56 likes786 downloads10mo agoHugging Face19coderchen01 /MMSD2.0 MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System This is a copy of the dataset uploaded on Hugging Face for easy access. The original data comes from this work, which is an improvement upon a previous study. Usage from typing import TypedDict, cast import pytorch_lightning as pl from datasets import Dataset, load_dataset from torch import Tensor from torch.utils.data import DataLoader from transformers import CLIPProcessor class… See the full description on the dataset page: https://huggingface.co/datasets/coderchen01/MMSD2.0.imagefeature-extraction10K<n<100K7 likes735 downloads2y agoHugging Face20CaraJ /MMSearch MMSearch 🔥: Benchmarking the Potential of Large Models as Multi-modal Search Engines Official repository for the paper "MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines". 🌟 For more details, please refer to the project page with dataset exploration and visualization tools: https://mmsearch.github.io/. [🌐 Webpage] [📖 Paper] [🤗 Huggingface Dataset] [🏆 Leaderboard] [🔍 Visualization] 💥 News [2024.09.25] 🌟 The evaluation code now… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MMSearch.imagequestion-answeringn<1K25 likes727 downloads6mo agoHugging Face21MMShmBlogs /hmblogs-v3text10M<n<100M0 likes511 downloads3y agoHugging Face22NCSOFT /K-MMStar K-MMStar We introduce K-MMStar, a Korean adaptation of the MMStar [1] designed for evaluating vision-language models. By translating the val subset of MMStar into Korean and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language. (We observe that there are unanswerable cases (e.g., multiple images required to answer the question but only has a single image, vague questions or options) in the… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-MMStar.image1K<n<10K12 likes429 downloads1y agoHugging Face23zli12321 /mmstarimage1K<n<10K0 likes420 downloads1y agoHugging Face24DjangoJungle /MMSearch-Plusimagen<1K0 likes386 downloads10mo agoHugging Face25amine-khelif /mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/mms_ulab_v2.audioaudio-to-audio10K<n<100K0 likes349 downloads7mo agoHugging Face26sensenova /HR-MMSearch Dataset Description HR-MMSearch is a benchmark designed to evaluate the Agentic Reasoning and Search capabilities of Multimodal Large Language Models in complex visual tasks. This dataset was introduced by SenseTime Research in the paper SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning. Key Features: High-Resolution Images: Contains high-resolution image inputs, requiring the model to possess fine-grained visual perception… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/HR-MMSearch.imagen<1K0 likes342 downloads9mo agoHugging Face27OmniSVG /MMSVGBench SVG Benchmark Dataset Dataset Description This dataset contains benchmark data for SVG generation tasks. Splits image2svg: Image to SVG conversion task (300 samples) text2svg: Text to SVG generation task (300 samples) Features Feature Type Description id string MD5 hash of the input (image bytes or text) image image Input image for image2svg task (None for text2svg) text string Input text for text2svg task (empty for… See the full description on the dataset page: https://huggingface.co/datasets/OmniSVG/MMSVGBench.imagen<1K7 likes340 downloads10mo agoHugging Face28jyjyjyjy /MMS-e MMS-e: Benchmarking the Resilience of Large Multimodal Models to Visual Scrambling Benchmark Examples Patchwise Question Answering: Divide the images into 2x2, 4x4, and 8x8 patches, then shuffle all the patches, and measure the ability of LMMs to answer questions about these images. Reconstruction task: Let LMMs reconstruct the order of shuffled patches based on the image' s caption, and let LMMs reconstruct the shuffled caption based on the image. Fixed Patch… See the full description on the dataset page: https://huggingface.co/datasets/jyjyjyjy/MMS-e.imagequestion-answering1K<n<10K0 likes307 downloads2y agoHugging Face29wangzn2001 /MM-SafetyBenchimage10K<n<100K0 likes306 downloads2y agoHugging Face30zhangkangning /mmskills Multimodal Skill Packages for General Visual Agents 515 public skill packages across Ubuntu desktop, macOS, Minecraft, and Mario environments. Overview | Preview | Contents | Download | Format | Statistics | Citation Overview This Hugging Face repository hosts the public MMSkills data packages: reusable multimodal procedural skills for visual agents. Each skill combines: a concise SKILL.md procedure; runtime_state_cards.json with state-matching… See the full description on the dataset page: https://huggingface.co/datasets/zhangkangning/mmskills.imagevisual-question-answeringn<1K4 likes301 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.