Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes97k downloads3y agoHugging Face02mlfoundations /dcvlm-baseline-200b DCVLM-Baseline (200B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool. A smaller 6.25B-token version is also available. ⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.imageimage-text-to-text10K<n<100K8 likes42k downloads3mo agoHugging Face03mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes37k downloads3y agoHugging Face04mlfoundations /dcvlm-balanced-200b DCVLM-Balanced (200B tokens) DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool. The instruction-heavy counterpart (DCVLM-baseline) is available as dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.imageimage-text-to-text10K<n<100K1 likes21k downloads2mo agoHugging Face05mlfoundations /datacomp_1b DataComp-1B This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.image1B<n<10B53 likes18k downloads3y agoHugging Face06mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face07jablonkagroup /chempile-mlift ChemPile-MLIFT A comprehensive multimodal dataset for chemistry property prediction using vision large language models 📋 Dataset Summary ChemPile-MLIFT is a dataset designed for multimodal chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using vision large language models (VLLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-mlift.imagetext-generation10M<n<100M14 likes15k downloads1y agoHugging Face08inductionlabs /frontier-ml-tasks Frontier MLE tasks Data for the tasks in induction-labs/frontier-ml. Each task has one folder: <slug>/public/ given to the coding agent verbatim, at /task/public <slug>/private/ held-out data for the trusted verifier, at /tests/private <slug>/artifacts/ optional starting files for the agent, at /artifacts <slug>/manifest.json optional provenance summary Tasks pin this repository by commit in tasks/<slug>/task.yaml. Starting model weights are downloaded from their… See the full description on the dataset page: https://huggingface.co/datasets/inductionlabs/frontier-ml-tasks.imagen<1K0 likes8.8k downloads18h agoHugging Face09mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes8.7k downloads2y agoHugging Face10mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes6.7k downloads2y agoHugging Face11GD-ML /MAPBench-V2For more details, please check our project page. Paper: https://arxiv.org/abs/2601.05432 Repository: https://github.com/AMAP-ML/Thinking-with-Map image1K<n<10K4 likes6.5k downloads8mo agoHugging Face12mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes6.4k downloads2y agoHugging Face13mlfoundations /DataComp-12M Dataset Card for DataComp-12M This dataset contains a 12M subset of DataComp-1B-BestPool. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Image-text models trained on DataComp-12M are significantly better than on CC-12M/YFCC-15M as well as DataComp-Small/Medium. DataComp-12M was introduced in MobileCLIP paper and along with the reinforced dataset DataCompDR-12M. The UIDs… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/DataComp-12M.imagetext-to-image14 likes4.3k downloads2y agoHugging Face14mlfoundations /dcvlm-baseline-6_25b DCVLM-Baseline (6.25B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This dataset version is a small 6.25B-token (small-pool) release consisting of 3,253,356 samples. ⚠️ NOTE: The training data is the WebDataset shards under shards/. The preview config shown in the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-6_25b.imageimage-text-to-textn<1K1 likes4.2k downloads3mo agoHugging Face15ifx-pse-sys-ml /FineVisionConcatShuffleIFXimage10M<n<100M0 likes4k downloads8mo agoHugging Face16mlfoundations-cua-dev /osworld-trajectoriesimage10K<n<100K0 likes3.5k downloads11mo agoHugging Face17mlfoundations /gelato-osworld-agent-trajectoriesimage10K<n<100K2 likes3.3k downloads11mo agoHugging Face18zr-zhang /MLLM-Generated-Image-Detection-Dataset MLLM-Generated Image Dataset This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection. Paper | Code Dataset Summary We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.imageimage-classification10K<n<100K1 likes3.1k downloads4d agoHugging Face19MLCommons /speech-wikimedia Dataset Card for Speech Wikimedia Dataset Summary The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers. Each audiofile should have one or more transcriptions in different languages. Transcription languages English German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.audion<1K14 likes2.7k downloads3y agoHugging Face20MLL-Lab /LAGEN-datasets LAGEN datasets Project resources: LAGEN collection. Training data, episode-level evaluations and paper experiment manifests. Experiment Entry Sim2Real calibration 30-case held-out calibration Visual history Visual history Latency in prompt Latency in prompt Latency transfer Latency transfer Task transfer Task transfer VLA fine-tuning scope VLA fine-tuning scope Mean vs. profile training Mean vs. profile training Observation stride Observation stride… See the full description on the dataset page: https://huggingface.co/datasets/MLL-Lab/LAGEN-datasets.imagerobotics10M<n<100M0 likes2.7k downloads2d agoHugging Face21ML-Intern-lab /gso-orbit-rgbaimagen<1K0 likes2.6k downloads16d agoHugging Face22mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes2.5k downloads2y agoHugging Face23EchoSafe-MLLM /MM-SafetyBench-plus-plus MM-SafetyBench++ Project Page | Paper | Code MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent. Dataset Summary For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.imageimage-text-to-text1K<n<10K2 likes1.8k downloads7mo agoHugging Face24MLL-Lab /MultiBBQ-perturbations MultiBBQ: image perturbations Image-level perturbation sets used for the robustness experiments in Fairness Failure Modes of Multimodal LLMs. Each set is the GPT-Image-1 image collection from MLL-Lab/MultiBBQ with a single, controlled transform applied. Evaluating on a perturbed set measures how stable a model's fairness behavior is under everyday image degradations. Paper: Fairness Failure Modes of Multimodal LLMs Code:… See the full description on the dataset page: https://huggingface.co/datasets/MLL-Lab/MultiBBQ-perturbations.imagevisual-question-answering1K<n<10K0 likes1.6k downloads15d agoHugging Face25mlfoundations /datacomp_medium DataComp Medium Pool This repository contains metadata files for the medium pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_medium.image100M<n<1B3 likes1.4k downloads3y agoHugging Face26mlfoundations-cua-dev /easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPimage10K<n<100K1 likes1.3k downloads1y agoHugging Face27mlfoundations /datacomp_small DataComp Small Pool This repository contains metadata files for the small pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_small.image10M<n<100M6 likes1.2k downloads3y agoHugging Face28MLL-Lab /MindCube MindCube: Spatial Mental Modeling from Limited Views MindCube is a novel benchmark designed to evaluate how well Vision Language Models (VLMs) can form robust spatial mental models from limited visual views. It comprises 21,154 questions across 3,268 images, assessing capabilities such as cognitive mapping (representing positions), perspective-taking (orientations), and mental simulation (dynamics for "what-if" movements). The dataset aims to expose critical gaps in existing VLMs'… See the full description on the dataset page: https://huggingface.co/datasets/MLL-Lab/MindCube.imagequestion-answering1K<n<10K11 likes1.2k downloads11mo agoHugging Face29mlfoundations /VisIT-Bench Dataset Card for VisIT-Bench Dataset Description Links Dataset Structure Data Fields Data Splits Data Loading Licensing Information Annotations Considerations for Using the Data Citation Information Dataset Description VisIT-Bench is a dataset and benchmark for vision-and-language instruction following. The dataset is comprised of image-instruction pairs and corresponding example outputs, spanning a wide range of tasks, from simple object recognition to complex… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/VisIT-Bench.imagen<1K16 likes1.1k downloads3y agoHugging Face30jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes974 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.