Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes5.4k downloads1mo agoHugging Face02Logics-MLLM /Logics-STEM-SFT-Dataset-Open-1.6M Logics-STEM-SFT-Dataset-2.2M 📰 News [2026.01.05]🔥 Release of our Techinical Report. [2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M. Overview What is this dataset? Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.text1M<n<10M33 likes2.1k downloads9mo agoHugging Face03EchoSafe-MLLM /MM-SafetyBench-plus-plus MM-SafetyBench++ Project Page | Paper | Code MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent. Dataset Summary For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.imageimage-text-to-text1K<n<10K2 likes1.8k downloads7mo agoHugging Face04Fancy-MLLM /R1-Onevision R1-Onevision [📂 GitHub][📝 Paper] [🤗 Reasoning Benchmark] [🤗 HF Demo] R1-Onevision Dataset Dataset Overview The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities. Aimed at bridging the gap between visual and textual understanding, this dataset provides rich, context-aware reasoning tasks across diverse domains, including natural scenes, science, mathematical problems… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision.textquestion-answering100K<n<1M47 likes1.6k downloads2y agoHugging Face05MLLM-CL /CL-VISTA MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark ✨Introduction • 🥇 Methods Provided • 🏦 Benchmarks • 🎨 Models 🏃 How to run • 🤝 Acknowledgments • 🙂 Contact If you like our project, please give us a star ⭐ on GitHub for the latest updates. ✨ Introduction MCITlib is a unified library for continual instruction tuning of multimodal large language models (MLLMs). It integrates diverse continual learning methods into a… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/CL-VISTA.text100K<n<1M1 likes1.3k downloads5mo agoHugging Face06qinbright99 /mllm-evaltext100K<n<1M0 likes1.1k downloads1y agoHugging Face07mllmTeam /DroidCall DroidCall: A Dataset for LLM-powered Android Intent Invocation paper|github DroidCall is the first open-sourced, high-quality dataset designed for fine-tuning LLMs for accurate intent invocation on Android devices. This repo contains data generated by DroidCall. The process of data generation is shown in the figure below Details can be found in our paper and github repository. What is Android Intent Invocation? Android Intent is a key machanism in Android that allows… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/DroidCall.texttext-generation10K<n<100K3 likes879 downloads2y agoHugging Face08MLLMMU /MLLMU-Bench Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench Abstract Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/MLLMMU/MLLMU-Bench.image1K<n<10K6 likes805 downloads2y agoHugging Face09MLLM-CL /UCITUnofficial training-ready fork of HaiyangGuo/UCIT image100K<n<1M1 likes735 downloads6mo agoHugging Face10Logics-MLLM /Logics-SWE-Env-2.5K Logics-SWE-Env-2.5K 2,553 software engineering task instances · 1,771 repositories · 4 programming languages 🤗 Related model: Logics-SWE-Qwen3.6-27B 📄 Paper: One to More, More to One 💻 GitHub: AgenticBigBang Overview What is this dataset? Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.tabulartext-generation1K<n<10K5 likes708 downloads18d agoHugging Face11MLLM-CL /Domain40kimage100K<n<1M1 likes437 downloads6mo agoHugging Face12Sarab-MLLMs /sarab Sarab Dataset The dataset behind Sarab, a cause-diagnostic Arabic visual hallucination evaluation benchmark for multimodal LLMs, modeled on Liu et al.'s CVPR 2025 PhD benchmark. Code and evaluation scripts are on GitHub. One real item from each mode, with the question in Arabic and English, the expected answer, and an actual model reply (a tick marks a correct first word, a cross a wrong one). What this is A human-captioned pool of Arabic Cultural Visual… See the full description on the dataset page: https://huggingface.co/datasets/Sarab-MLLMs/sarab.imagevisual-question-answeringn<1K0 likes429 downloads38m agoHugging Face13PaDT-MLLM /RefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.textobject-detection100K<n<1M4 likes394 downloads1y agoHugging Face14Logics-MLLM /OmniParsingBench 🤗 Model   |   📑 Technical Report   |   💻 GitHub OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities. Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.image1K<n<10K2 likes393 downloads6mo agoHugging Face15lingshu-medical-mllm /ReasonMed ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning 📄 Paper  |  💻 Code  |  📊 Dataset ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.textquestion-answering1M<n<10M95 likes358 downloads1y agoHugging Face16Logics-MLLM /Logics-STEM-SFT-Dataset-Open-5.3Mtext1M<n<10M4 likes347 downloads9mo agoHugging Face17MLLM-CL /VTCBench Dataset Card for VTCBench Vision-Text Compression Benchmark (VTCBench) revisits Needle-In-A-Haystack (NIAH) from a VLM's perspective by converting long context into rendered images. This benchmark tests VLM's ability to OCR, retrieve, aggregate, infer, and memorize long context as images. Specifically, this benchmark includes 3 tasks: Retrieval: Vision-NIAH VQA task for information retrieval and aggregation.… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/VTCBench.imagevisual-question-answering1K<n<10K4 likes251 downloads2mo agoHugging Face18MLLM-CL /DCLimage100K<n<1M0 likes201 downloads5mo agoHugging Face19Juancc-ctic /MLLM_agriculture_evaluation Evaluation methodology dataset for multimodal LLM-based system with geospatial augmented context for agricultural plot-level assistance This dataset was built for evaluating context-amplified multimodal Large Language Models (LLM) within the agricultural domain. It is supplementary material for a manuscript currently under peer review. Data Structure The dataset contains the following columns visible in the viewer table above: QueryID (string): The unique… See the full description on the dataset page: https://huggingface.co/datasets/Juancc-ctic/MLLM_agriculture_evaluation.imagen<1K0 likes197 downloads15d agoHugging Face20Pawlo77 /mllm-shap MLLM-SHAP experiment datasets Curated test splits for studying Shapley-value explanations in multimodal large language models (text and audio inputs). Each configuration is a filtered, size-controlled subset built for reproducible benchmarking—not a full copy of the upstream corpora. Configs follow the naming pattern {task}__{source} (for example single_sentence__voice_bench). Quick load Pin a dataset revision for reproducibility (replace REVISION with the commit hash… See the full description on the dataset page: https://huggingface.co/datasets/Pawlo77/mllm-shap.tabulartext-generation1K<n<10K2 likes182 downloads5mo agoHugging Face21VITA-MLLM /Comic-9K Comic-9K Image Extracting all images. cat images.tar.gz.aa images.tar.gz.ab images.tar.gz.ac images.tar.gz.ad images.tar.gz.ae > images.tar.gz tar xvzf images.tar.gz Summary We provide human-written plot synopsis. summary.jsonl image100K<n<1M6 likes156 downloads2y agoHugging Face22MLLM-CL /MLLM-CL-ReplayData MLLM-CL: Continual Learning for Multimodal Large Language Models This is the official dataset repository of MLLM-CL and MR-LoRA. MLLM-CL is a novel benchmark encompassing domain and ability continual learning, where the former focuses on independently and identically distributed (IID) evaluation across evolving mainstream domains, whereas the latter evaluates on non-IID scenarios with emerging model ability. MR-LoRA prevents catastrophic interference through parameter isolation and… See the full description on the dataset page: https://huggingface.co/datasets/MLLM-CL/MLLM-CL-ReplayData.textimage-text-to-text100K<n<1M0 likes147 downloads1y agoHugging Face23PaDT-MLLM /COCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/COCO.textobject-detection100K<n<1M1 likes147 downloads1y agoHugging Face24zhhxte /mllm_cl_vizwizimage10K<n<100K0 likes115 downloads1y agoHugging Face25Frankdz /MLLM-CL-RAGtext10K<n<100K0 likes110 downloads6mo agoHugging Face26zhhxte /mllm_cl_textvqaimage10K<n<100K0 likes109 downloads1y agoHugging Face27tomyoon2 /MLLM_testimage10K<n<100K0 likes102 downloads11mo agoHugging Face28Fancy-MLLM /R1-Onevision-Bench R1-Onevision-Bench [📂 GitHub][📝 Paper] [🤗 HF Dataset] [🤗 HF Model] [🤗 HF Demo] Dataset Overview R1-Onevision-Bench comprises 38 subcategories organized into 5 major domains, including Math, Biology, Chemistry, Physics, Deducation. Additionally, the tasks are categorized into five levels of difficulty, ranging from ‘Junior High School’ to ‘Social Test’ challenges, ensuring a comprehensive evaluation of model capabilities across varying complexities.… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision-Bench.textquestion-answeringn<1K3 likes98 downloads2y agoHugging Face29PaDT-MLLM /ReferringImageCaptioningPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/ReferringImageCaptioning.textimage-to-text100K<n<1M3 likes98 downloads1y agoHugging Face30Lris47 /MLLM3-textcaps-scienceqa-vqav2image100K<n<1M0 likes86 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.