Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ComplexDataLab /OpenFake Dataset Card for OpenFake OpenFake is a dataset and benchmark for detecting AI-generated images, with a focus on politically and socially salient content where misinformation risk is highest. It pairs real photographs with synthetic counterparts produced by a wide range of frontier proprietary generators, open-source diffusion models, and community fine-tunes. A separate in-the-wild test set is sourced from Reddit to evaluate detector performance on naturally circulated… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.imageimage-classification1M<n<10M35 likes17k downloads11d agoHugging Face02nebula /OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World. Code: https://github.com/iamwangyabin/OpenSDI imageimage-classification100K<n<1M2 likes4.2k downloads2y agoHugging Face03nebula /OpenSDI_test OpenSDI: Spotting Diffusion-Generated Images in the Open World This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper: Project Page: https://iamwangyabin.github.io/OpenSDI/ OpenSDID Dataset Highlights: User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs. Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.imageimage-classification100K<n<1M1 likes3k downloads2y agoHugging Face04LAYEK-143 /Open-Pixel-1T 🌌 Open-Pixel-1T (Visual Atlas) A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training 📑 Dataset Summary Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.imagetext-to-image10M<n<100M8 likes2.7k downloads6mo agoHugging Face05jaddai /openbrush OpenBrush-75K A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research. Dataset Description OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush.tabularimage-to-text10K<n<100K3 likes2k downloads16d agoHugging Face06MichalMlodawski /closed-open-eyes 👀 Open and Closed Eyes Dataset Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟 📁 Dataset Structure The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/MichalMlodawski/closed-open-eyes.imageimage-classification100K<n<1M5 likes1.5k downloads2y agoHugging Face07nyuuzyou /OpenGameArt-CC0 Dataset Card for OpenGameArt-CC0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.audioimage-classification10K<n<100K10 likes1.1k downloads1y agoHugging Face08nyuuzyou /openclipart Dataset Card for OpenClipart.org SVG Images Dataset Summary This dataset contains 178,604 public domain SVG vector clipart images collected from OpenClipart.org. OpenClipart.org is a community-driven platform where artists share vector clip art explicitly released into the public domain (CC0). The dataset includes the SVG content along with comprehensive metadata such as titles, descriptions, artist names, creation dates, tags, and image URLs. The SVG files in this… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/openclipart.textimage-classification100K<n<1M9 likes789 downloads1y agoHugging Face09Rapidata /OpenAI-4o_t2i_human_preference Rapidata OpenAI 4o Preference This T2I dataset contains over 200'000 human responses from over ~45,000 individual annotators, collected in less than half a day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenAI-4o_t2i_human_preference.imagetext-to-image10K<n<100K34 likes586 downloads2y agoHugging Face10Rapidata /OpenGVLab_Lumina_t2i_human_preference Rapidata Lumina Preference This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Lumina across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenGVLab_Lumina_t2i_human_preference.imagetext-to-image10K<n<100K13 likes476 downloads2y agoHugging Face11nebula /OpenSDIDplus OpenSDID+ OpenSDID+ is an extended release of the OpenSDI dataset. It complements the original SD1.5 training split with large-scale images from the remaining OpenSDI generators: SD2, SD3, SDXL, and FLUX. The dataset follows the OpenSDI challenge introduced in "OpenSDI: Spotting Diffusion-Generated Images in the Open World". OpenSDI studies detection and localization of diffusion-generated images under realistic open-world settings, including diverse user intentions, evolving… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDIDplus.imageimage-classification100K<n<1M1 likes443 downloads4mo agoHugging Face12timm /imet-open-access iMet Open Access A multi-label image classification dataset of 259,559 artworks and objects from The Metropolitan Museum of Art, built from the Met's official CC0 Open Access release (metmuseum/openaccess). It follows the spirit of the Kaggle iMet Collection challenges (FGVC6 2019 / FGVC7 2020): long-tail, fine-grained attribute recognition from museum photography, with subject tags, culture, medium (material / technique) and country labels. This is not the Kaggle iMet data.… See the full description on the dataset page: https://huggingface.co/datasets/timm/imet-open-access.imageimage-classification100K<n<1M1 likes404 downloads12d agoHugging Face13jaddai /openbrush-baroque OpenBrush Baroque Baroque works from OpenBrush-75K (~1600–1750) — chiaroscuro, religious painting, dramatic light. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,240 you actually want. Why this subset The canonical Baroque visual language — Caravaggio, Rembrandt, Vermeer, Velázquez, Rubens. Useful for models learning dramatic… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-baroque.tabularimage-to-text1K<n<10K1 likes349 downloads16d agoHugging Face14jaddai /openbrush-impressionism OpenBrush Impressionism Every Impressionist work from OpenBrush-75K — the largest movement subset. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,798 you actually want. Why this subset Broad-coverage subset for training on the Impressionist visual language: broken brushwork, light-on-color theory, plein-air staging, atmospheric… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionism.tabularimage-to-text10K<n<100K1 likes331 downloads16d agoHugging Face15Trever896 /openbrush-75k OpenBrush-75K A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research. Dataset Description OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a vision-language… See the full description on the dataset page: https://huggingface.co/datasets/Trever896/openbrush-75k.tabularimage-to-text10K<n<100K1 likes288 downloads7mo agoHugging Face16opendiffusionai /cc12m-cleaned CC12m-cleaned This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by CaptionEmporium (The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_ I have then used the llava captions as a base, and used the detailed descrptions to filter out images with things like watermarks, artist signatures, etc. I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.imagetext-to-image1M<n<10M13 likes285 downloads2y agoHugging Face17jaddai /openbrush-landscapes OpenBrush Landscapes Every landscape painting from OpenBrush-75K — across all artists, movements, and centuries. Largest single-genre subset. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,612 you actually want. Why this subset Every landscape across the parent dataset's full range — Romantic wildernesses, Impressionist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-landscapes.tabularimage-to-text10K<n<100K1 likes260 downloads16d agoHugging Face18IamBusy /OpenJev-Vision-Research-v0.1 OpenJev Vision Research v0.1 12,832 image records, with public provenance, original synthetic scenes, and programmatically derived decision questions. This is an experimental research dataset for visual posterior learning and compositional decisions, released with OpenJev. It is not a reproduction of TypeSafe's proprietary Jev model or training method. Three separate configurations Config Images What the labels mean License synthetic 8,192 Exact… See the full description on the dataset page: https://huggingface.co/datasets/IamBusy/OpenJev-Vision-Research-v0.1.imageimage-classification10K<n<100K0 likes244 downloads21d agoHugging Face19V4ldeLund /nationalmuseet-open-images Nationalmuseet Open Images This dataset is an independently harvested research dataset from Nationalmuseet Samlinger Online. It contains metadata and optionally WebDataset image shards for Nationalmuseet asset records whose rights.license is one of: Public Domain CC-BY No known rights Public Domain and CC-BY are the strict open-license subset. No known rights is kept as a separate license bucket because Nationalmuseet says this label means that, to their best assessment, the… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/nationalmuseet-open-images.imageimage-classification100K<n<1M0 likes214 downloads5mo agoHugging Face20VasilyLoginov /closed-open-eyes 👀 Open and Closed Eyes Dataset Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟 📁 Dataset Structure The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/VasilyLoginov/closed-open-eyes.imageimage-classification100K<n<1M1 likes197 downloads10mo agoHugging Face21nyuuzyou /OpenGameArt-CC-BY-3.0 Dataset Card for OpenGameArt-CC-BY-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 3.0 (CC-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-3.0.textimage-classification1K<n<10K1 likes195 downloads1y agoHugging Face22nyuuzyou /OpenGameArt-CC-BY-SA-3.0 Dataset Card for OpenGameArt-CC-BY-SA-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution-ShareAlike 3.0 Unported (CC-BY-SA-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-SA-3.0.audioimage-classification1K<n<10K1 likes192 downloads1y agoHugging Face23nyuuzyou /OpenGameArt-OGA-BY-3.0 Dataset Card for OpenGameArt-OGA-BY-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution (OGA-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions and metadata are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-3.0.audioimage-classificationn<1K1 likes183 downloads1y agoHugging Face24jaddai /openbrush-impressionist-landscapes OpenBrush Impressionist Landscapes Cross-cut subset: Impressionist landscape paintings from OpenBrush-75K. The most-targeted style+genre combination for Impressionist landscape LoRA training. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,308 you actually want. Why this subset The intersection of the largest movement… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionist-landscapes.tabularimage-to-text1K<n<10K1 likes179 downloads16d agoHugging Face25jaddai /openart-items-artifacts OpenArt — Items & Artifacts openart-items-artifacts is the items artifacts subject collection of the OpenArt family of open, public-domain art datasets: 25,750 works (11,317 paintings/illustrations · 14,216 photographed objects · 217 unclassified), each paired with a structured VLM caption plus medium, attribution and inscription metadata. Human-made objects and the decorative arts — vessels, tools, arms and armor, textiles, furniture and ornament — both as physical artifacts… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-items-artifacts.tabularimage-to-text10K<n<100K2 likes174 downloads4mo agoHugging Face26nyuuzyou /OpenGameArt-Mixed-Licenses Dataset Card for OpenGameArt-Mixed-Licenses Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.audioimage-classification1K<n<10K0 likes171 downloads1y agoHugging Face27jaddai /openart-portraits-classical OpenArt — Portraits & the Classical Figure openart-portraits-classical is the portraits classical subject collection of the OpenArt family of open, public-domain art datasets: 28,011 works (13,868 paintings/illustrations · 13,970 photographed objects · 173 unclassified), each paired with a structured VLM caption plus medium, attribution and inscription metadata. The human figure and portraiture across the full range of media — painted and drawn portraits alongside photographic… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-portraits-classical.tabularimage-to-text10K<n<100K1 likes155 downloads4mo agoHugging Face28jaddai /openbrush-religious-art OpenBrush Religious Art Religious paintings from OpenBrush-75K — saints, biblical scenes, devotional works. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,119 you actually want. Why this subset A coherent visual genre: religious narrative painting from medieval through early modern. Heavy on Renaissance and Baroque eras. Common… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-religious-art.tabularimage-to-text1K<n<10K1 likes152 downloads16d agoHugging Face29openpecha /uchen_ume_classification_dataset Uchen–Ume Classification Benchmark A binary image classification dataset for distinguishing two fundamental categories of Tibetan script: Uchen (དབུ་ཅན།, headed script with a horizontal top stroke) and Ume (དབུ་མེད།, headless script without a top stroke). All images are raw, unprocessed manuscript scans from the Buddhist Digital Resource Center (BDRC). Model: openpecha/uchen-ume-classifier Dataset summary Split Examples Uchen Ume Train 9,110 ~3,124 ~5,986… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/uchen_ume_classification_dataset.imageimage-classification10K<n<100K0 likes140 downloads5mo agoHugging Face30nyuuzyou /OpenGameArt-CC-BY-4.0 Dataset Card for OpenGameArt-CC-BY-4.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual: English… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-4.0.audioimage-classification1K<n<10K2 likes135 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.