Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes15k downloads4mo agoHugging Face02biglam /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.imageimage-classification1M<n<10M64 likes6.3k downloads2mo agoHugging Face03suvadityamuk /amazon-berkeley-objects Amazon Berkeley Objects (ABO) A Hugging Face packaging of the Amazon Berkeley Objects (ABO) dataset. The data content is the official CC BY 4.0 release from https://amazon-berkeley-objects.s3.amazonaws.com/index.html. This mirror changes only the packaging: files are grouped into typed Parquet shards, and every original media file is preserved byte-for-byte and never transcoded. Images use the datasets Image() feature, 3D product models use the native Mesh() feature (original… See the full description on the dataset page: https://huggingface.co/datasets/suvadityamuk/amazon-berkeley-objects.imageimage-classification1M<n<10M6 likes3.8k downloads2mo agoHugging Face04jaddai /openbrush OpenBrush-75K A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research. Dataset Description OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush.tabularimage-to-text10K<n<100K3 likes3k downloads11d agoHugging Face05Faizaniqbal /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.imageimage-classification1M<n<10M0 likes3k downloads1mo agoHugging Face06sewa-rural-care /anemia-survey-datasetgated Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India Dataset: sewa-rural-care/anemia-survey-dataset Contact: sewarural@ymail.com Version: 1.0 — July 2026 Dataset Summary This dataset supports research into non-invasive, smartphone-based anemia screening applicable to low-resource and rural healthcare settings. It was collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.tabularimage-classification1K<n<10K7 likes1.3k downloads3mo agoHugging Face07trojblue /danbooru2025-metadata 🎨 Danbooru 2025 Metadata Latest Post ID: 9,158,800 (as of Apr 16, 2025) 📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork. Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes: More consistent tag history tracking Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.imagetext-to-image1M<n<10M38 likes1.2k downloads1y agoHugging Face08purvanshi /TASTE TASTE: Human Preferences for Design-Quality Image Comparison This dataset is the human-evaluation corpus released alongside the TASTE preference model. It contains panel rankings of generated images across multiple quality dimensions — both aesthetic (does the image look good?) and description-faithfulness (does the image match what the prompt describes?) — plus a per-image hallucination judgement. Quick stats Table Rows Notes prompts.parquet ~200 one… See the full description on the dataset page: https://huggingface.co/datasets/purvanshi/TASTE.imageimage-classification10K<n<100K9 likes1.2k downloads4mo agoHugging Face091aurent /PovertyMap PovertyMap-Wilds: Poverty mapping across different countries Description This is a processed version of LandSat 5/7/8 satellite imagery originally from Google Earth Engine under the names LANDSAT/LC08/C01/T1_SR,LANDSAT/LE07/C01/T1_SR,LANDSAT/LT05/C01/T1_SR, nighttime light imagery from the DMSP and VIIRS satellites (Google Earth Engine names NOAA/DMSP-OLS/CALIBRATED_LIGHTS_V4 and NOAA/VIIRS/DNB/MONTHLY_V1/VCMSLCFG) and processed DHS survey metadata obtained from… See the full description on the dataset page: https://huggingface.co/datasets/1aurent/PovertyMap.tabularimage-classification10K<n<100K1 likes1.1k downloads2y agoHugging Face10andropar /relaion2b-natural-embeddings LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart) LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7). Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.tabularfeature-extraction100M<n<1B1 likes1.1k downloads6mo agoHugging Face11biglam /britannica-illustrated-pages Britannica Illustrated Pages 115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition (1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes (838 Internet Archive items). A second config carries the classifier score, OCR word count and provenance for every one of the 975,345 pages. Two things the scan showed: 82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.imageimage-classification1M<n<10M49 likes1.1k downloads1mo agoHugging Face12bazyl /GTSRB Dataset Card for GTSRB Dataset Summary The German Traffic Sign Benchmark is a multi-class, single-image classification challenge held at the International Joint Conference on Neural Networks (IJCNN) 2011. We cordially invite researchers from relevant fields to participate: The competition is designed to allow for participation without special domain knowledge. Our benchmark has the following properties: Single-image, multi-class classification problem More than 40… See the full description on the dataset page: https://huggingface.co/datasets/bazyl/GTSRB.tabularimage-classification10K<n<100K0 likes1.1k downloads4y agoHugging Face13debajyotidasgupta /vecforge-paper-corpus Note (rebuild in progress): figures are being re-extracted with a fixed extractor (cleaner crops). The image-preview config (Parquet with an inline column + difficulty/type/score labels) returns after re-classification. The config (paper metadata + links) is live now. VecForge Paper Corpus A pristine, deduplicated collection of 59,732 top-venue AI/ML/CV/NLP/Robotics papers (2020-2024) with every captioned figure and full paper text, for research on figure understanding… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/vecforge-paper-corpus.imageimage-to-text10K<n<100K0 likes945 downloads4mo agoHugging Face14Ardea /Icarus-dataset Icarus A unified multi-modal curriculum dataset for evolutionary neural architecture search. Every row is one self-contained Task = {meta, support, query}, where support and query are lists of (input_Field, output_Field) pairs. The inner loop trains on support; fitness is scored on query. Support is non-empty for every task. Encoders read the Field descriptor (axes, value_type, n_classes, value_range, mask); mask is True where a value is padding/ignored. meta.class_names, when… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/Icarus-dataset.textimage-classification10K<n<100K2 likes915 downloads4mo agoHugging Face15Waheed786dar /Comiman-Dataset Comiman Dataset Attribution is required for every use: Comiman Dataset by Waheed (huggingface.co/Waheed786dar) - https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset A license-gated comics and manga page corpus built for training a model that can plan and draw full comic/manga series (the planned model: Waheed786dar/Comiman). Every book passed an automatic license gate (Creative Commons / CC0 / Public Domain Mark metadata, or a public-domain claim limited to works… See the full description on the dataset page: https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset.tabularimage-to-text10K<n<100K0 likes889 downloads3d agoHugging Face16yxma /gelsight-mini-pretrain GelSight Mini Pretrain ~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92. Frames Sources Real 536K FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad Sim 317K sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated) NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.imageimage-classification100K<n<1M0 likes876 downloads2mo agoHugging Face17vimageiitb /GeoMeld 🌍 GeoMeld Multi-Modal Earth Observation Dataset (WebDataset) GeoMeld is a large-scale multi-modal remote sensing dataset introduced in our CVPRW 2026 paper on semantically grounded foundation modeling. GeoMeld contains approximately 2.5 million spatially aligned samples spanning heterogeneous sensing modalities and spatial resolutions, paired with semantically grounded captions generated through an agentic pipeline. The dataset is designed to support multimodal representation… See the full description on the dataset page: https://huggingface.co/datasets/vimageiitb/GeoMeld.tabularimage-classificationn<1K3 likes812 downloads4mo agoHugging Face18mlech26l /liquidrandom-data liquidrandom-data Diverse seed data for ML/LLM training data generation pipelines. Used by the liquidrandom Python package. Dataset Summary This dataset contains 520,080 seed data samples across 24 categories, generated using a hierarchical taxonomy tree approach with LLM-based quality validation and fuzzy deduplication. Data is stored as Parquet with zstd compression. Categories Category Samples File Coding Tasks 30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.tabulartext-generation100K<n<1M0 likes808 downloads2mo agoHugging Face19do-me /Flickr-Geo Dataset Card for Flickr-Geo 217.646.487 images with lat lon coordinates from Flickr. This repo is a filtered version of https://huggingface.co/datasets/bigdata-pw/Flickr for all rows containing a valid lat lon pair. Load and Visualize Load the data (2s on my system) with DuckDB: import duckdb df = duckdb.sql(""" SELECT CAST(latitude AS DOUBLE) AS latitude, CAST(longitude AS DOUBLE) AS longitude FROM 'reduced_flickr_data/*.parquet' """).df() df… See the full description on the dataset page: https://huggingface.co/datasets/do-me/Flickr-Geo.tabularimage-classification100M<n<1B26 likes694 downloads2y agoHugging Face20kakaobrain /coyo-labeled-300m Dataset Card for COYO-Labeled-300M Dataset Summary COYO-Labeled-300M is a dataset of machine-labeled 300M images-multi-label pairs. We labeled subset of COYO-700M with a large model (efficientnetv2-xl) trained on imagenet-21k. We followed the same evaluation pipeline as in efficientnet-v2. The labels are top 50 most likely labels out of 21,841 classes from imagenet-21k. The label probabilies are provided rather than label so that the user can select threshold of their… See the full description on the dataset page: https://huggingface.co/datasets/kakaobrain/coyo-labeled-300m.imageimage-classification100M<n<1B12 likes534 downloads4y agoHugging Face21jaddai /openbrush-landscapes OpenBrush Landscapes Every landscape painting from OpenBrush-75K — across all artists, movements, and centuries. Largest single-genre subset. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,612 you actually want. Why this subset Every landscape across the parent dataset's full range — Romantic wildernesses, Impressionist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-landscapes.tabularimage-to-text10K<n<100K1 likes495 downloads11d agoHugging Face22drksci /trade_vision_dataset TradeVision: Hierarchical Physical Business & Multimodal Retail Provenance Dataset This dataset is continuously seeded from OpenStreetMap, matched to Google Place IDs, harvested for temporal store photos, and enriched with zero-shot computer vision using Hugging Face Hub native pipelines. Dataset Structure The dataset is partitioned into three relational subsets loadable via Hugging Face datasets: from datasets import load_dataset # 1. Load Canonical Businesses… See the full description on the dataset page: https://huggingface.co/datasets/drksci/trade_vision_dataset.tabularzero-shot-object-detectionn<1K0 likes433 downloads23h agoHugging Face23gokhankocmarli /inline-digital-holography-v3 Dataset Card for Synthetic Inline Holographical Images v3 (224px Highly Diverse) This dataset provides synthetic image triplets representing inline holographical imaging in a simulated environment. This version (v3) uses a native 224x224 resolution optimized for modern Vision Transformers (ViT, Swin) and contains 25,000 samples across 8 noise configurations. Each data sample consists of: An object-domain field (ground truth), Its corresponding forward-propagated hologram (the… See the full description on the dataset page: https://huggingface.co/datasets/gokhankocmarli/inline-digital-holography-v3.tabularimage-to-image1B<n<10B0 likes418 downloads7mo agoHugging Face24deb0naire /Bodhisetu Bodhisetu A multimodal cultural-heritage dataset from India, collected via the SMRITI field-data-collection platform. Overview Bodhisetu documents 3,000+ heritage entities across 6 Indian languages (Tamil, Kannada, Assamese, Maithili, Bodo, Nepali) and 14 taluks spanning 5 states. Each entity is captured as one or more images and short-form videos, accompanied by field descriptions. This release contains 5,753 resources (4,461 images + 1,292 videos, ~91 GB)… See the full description on the dataset page: https://huggingface.co/datasets/deb0naire/Bodhisetu.imageimage-classification10K<n<100K0 likes408 downloads5mo agoHugging Face25liranmao /meowcat-predictions MeowCat cell-type predictions on TCGA-LUAD and CPTAC-CCRCC Per-pixel cell-type predictions generated by MeowCat on H&E whole-slide images from two public cohorts: Cohort Tissue Samples h5ad payload TCGA-LUAD Lung adenocarcinoma 531 ~60 GB CPTAC-CCRCC Clear-cell renal cell carcinoma 831 ~93 GB File layout composition.parquet # long format: sample × cell_type → count, fraction metadata.parquet # sample_id, cohort, patient_id, n_pixels… See the full description on the dataset page: https://huggingface.co/datasets/liranmao/meowcat-predictions.tabularimage-classification10K<n<100K0 likes399 downloads28d agoHugging Face26joelleoqiyi /trace-rx-eval-predictions TRACE-RX Evaluation Predictions Per-image detector scores from an independent evaluation of the two TechJam 2026 TRACE-RX detectors, run 30 Aug – 1 Sep 2026. No images here. Every file contains scores, labels, asset ids and transform names only — this is derived evaluation metadata, not a redistribution of any source imagery. The underlying corpora (Joshyxwa/data_draft, Joshyxwa/techjam2026, techjam-aigc/wildfake-eval-subset) keep their own terms, and data_draft's WildFake rows… See the full description on the dataset page: https://huggingface.co/datasets/joelleoqiyi/trace-rx-eval-predictions.tabularimage-classification100K<n<1M0 likes394 downloads1mo agoHugging Face27aiacademy-kg /house_kg_full_dataset house.kg — Kyrgyzstan Real Estate (multimodal) A complete snapshot of house.kg, the largest real-estate board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller identities, agency ratings, reviews — and 227,294 photographs. Field names are English; values are kept in the original language (Russian/Kyrgyz), exactly as the site renders them. 💻 Scraper source code on GitHub → The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.imagetabular-regression100K<n<1M0 likes371 downloads3mo agoHugging Face28ODELIA-AI /ODELIA-Challenge-2025gated ODELIA Challenge Dataset This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning. The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.tabularimage-classification1K<n<10K12 likes339 downloads1y agoHugging Face29Trever896 /openbrush-75k OpenBrush-75K A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research. Dataset Description OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a vision-language… See the full description on the dataset page: https://huggingface.co/datasets/Trever896/openbrush-75k.tabularimage-to-text10K<n<100K1 likes338 downloads7mo agoHugging Face30nesteo-datasets /nesteo-prototype NestEO: Modular and Hierarchical EO Dataset Framework NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO. Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.tabularimage-segmentation10K<n<100K1 likes333 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.