Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K55 likes18k downloads10mo agoHugging Face02MrigLabIITRopar /GroMo25 GroMo25: Multiview Time-Series Plant Image Dataset for Age Estimation and Leaf Counting Dataset Summary GroMo25 is a multiview, time-series plant image dataset designed for plant age estimation (in days) and leaf counting tasks in precision agriculture. It contains high-quality images of four crop species — Wheat, Okra, Radish, and Mustard — captured over multiple days under controlled conditions. Each plant is photographed from 24 angles across 5 vertical levels per day… See the full description on the dataset page: https://huggingface.co/datasets/MrigLabIITRopar/GroMo25.imageimage-classification100K<n<1M2 likes13k downloads6mo agoHugging Face03imageomics /fish-vista Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images. See Example Code to Use the Segmentation Dataset Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset. Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.imageimage-classification10K<n<100K26 likes13k downloads9mo agoHugging Face04vsevolodpl /REPID REPID: Rendering Evaluation of Photographic Image Dataset REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA) in paper Beyond distortions: a benchmark for subjective evaluation of image rendering quality. Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.imageimage-classification100K<n<1M4 likes4.8k downloads2mo agoHugging Face05anonymousllbench /llbench-dataset LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute. LL-Bench is a large-scale, human-preference benchmark for evaluating low-level vision restoration in the era of large generative models (LGMs). It compares 10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.imageimage-to-image100K<n<1M0 likes1.3k downloads5mo agoHugging Face06imageomics /questFish2024 Dataset Card for QUEST Fish 2024 Images collected by teachers during a QUEST workshop. In 2024, the images were of fish collected from bodies of water near Princeton University. Dataset Details Dataset Structure /dataset/ <folder>/ <img_id 1>.png <img_id 2>.png ... <img_id n>.png ... <img_id 1>.png <img_id 2>.png ... <img_id n>.png fieldData2024.csv Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/questFish2024.imageimage-classificationn<1K0 likes1.3k downloads2mo agoHugging Face07dartbrains /localizer Dartbrains Localizer Dataset A subset of the Brainomics/Localizer functional MRI dataset, prepared for the Dartbrains neuroimaging course at Dartmouth College. Quick Start Load beta maps (recommended for most exercises) from datasets import load_dataset ds = load_dataset("dartbrains/localizer", "betas") img = ds[0]["nifti"] # nibabel.Nifti1Image subject = ds[0]["subject"] # "S01" condition = ds[0]["condition"] # "audio_computation"… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/localizer.imageimage-classificationn<1K1 likes1.1k downloads3mo agoHugging Face08harvardairobotics /FairVision Dataset Card: Harvard-FairVision Dataset Summary Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes. This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.imageimage-classification10K<n<100K1 likes1k downloads6mo agoHugging Face09ramankamran /retina-age-analysis Retina Age Analysis Dataset Dataset Description This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks. Dataset Summary Task: Age prediction from retinal fundus images Images: 9,857 high-quality retinal images Patients: 5,393 unique patients Age Range: 5-97 years Image Format: JPEG Average Image Size: ~1 MB Supported Tasks Regression: Predict continuous age (5-97 years) Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.imageimage-classification1K<n<10K0 likes953 downloads1y agoHugging Face10AliHome3D /SA-BENCH SA-BENCH SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.” It evaluates the spatial aesthetics of interior images along four dimensions: distortion harmony layout lighting The dataset contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and evaluation. Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AliHome3D/SA-BENCH.imageimage-classification10K<n<100K0 likes803 downloads5mo agoHugging Face11waticlems /pcabiop PCaBiop PCaBiop is a whole-slide-image benchmark for measuring the robustness of slide-level pathology foundation models to non-biological variation. It pairs a controlled two-centre prostate-biopsy cohort, in which the contributing centre acts as a labelled confounder, with an external out-of-domain cohort for downstream shortcut-learning evaluation. The benchmark accompanies the paper A distributional robustness margin for pathology foundation models (citation below… See the full description on the dataset page: https://huggingface.co/datasets/waticlems/pcabiop.imageimage-classification1K<n<10K0 likes756 downloads2mo agoHugging Face12notgoodkeeper /cnn-based-drowsiness-detection-data CNN-Based Drowsiness Detection - Dataset Preprocessed, auto-labeled face-crop images used to train the model in notgoodkeeper/cnn-based-drowsiness-detection. Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection Collection Frames were captured from a webcam, then run through: Haar Cascade face detection -> crop + pad + resize to 412x412 MediaPipe Selfie Segmentation -> background replaced with white CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.imageimage-classification1K<n<10K1 likes752 downloads1mo agoHugging Face13SyedNazmusSakib /PlantInquiryVQA PlantInquiryVQA — Thinking Like a Botanist Benchmark and framework for multi-turn, intent-driven visual question answering in plant pathology. Accepted at ACL 2026 Findings. Overview PlantInquiryVQA formalises diagnostic reasoning in plant pathology as a Chain-of-Inquiry (CoI) — an ordered sequence of visually-grounded questions that adapts to the plant's severity and the expert's epistemic intent (Diagnosis / Prognosis / Management). The benchmark evaluates… See the full description on the dataset page: https://huggingface.co/datasets/SyedNazmusSakib/PlantInquiryVQA.imagevisual-question-answering100K<n<1M1 likes673 downloads5mo agoHugging Face14ykotseruba /SNAP SNAP Benchmark Code and annotations: [https://github.com/ykotseruba/SNAP] SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings. This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms. SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.imageimage-classification10K<n<100K0 likes624 downloads1y agoHugging Face15imageomics /Heliconius-Collection_Cambridge-Butterfly Dataset Card for Heliconius Collection (Cambridge Butterfly) Dataset Description Dataset Summary Subset of the collection records from Chris Jiggins' research group at the University of Cambridge, collection covers nearly 20 years of field studies. This subset contains approximately 36,189 RGB images of 11,962 specimens (29,134 images of 10,086 specimens across all Heliconius). Many records have both images and locality data. Most images were… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Heliconius-Collection_Cambridge-Butterfly.imageimage-classification10K<n<100K2 likes523 downloads1y agoHugging Face16Humanbased-AI /MM-Food-100K Overview This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/MM-Food-100K.imageimage-classification100K<n<1M79 likes428 downloads1y agoHugging Face17gaoyuan-ai /SA-BENCH SA-BENCH SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.” Accepted to CVPRW 2026. GitHub | CVF Open Access | arXiv | Model It evaluates the spatial aesthetics of interior images along four dimensions: distortion harmony layout lighting SA-BENCH contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and… See the full description on the dataset page: https://huggingface.co/datasets/gaoyuan-ai/SA-BENCH.imageimage-classification10K<n<100K5 likes400 downloads2mo agoHugging Face18yzhllm /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K0 likes394 downloads5mo agoHugging Face19harvardairobotics /FairVLMed Dataset Card: Harvard-FairVLMed Dataset Summary Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language. This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.imageimage-classification10K<n<100K0 likes363 downloads6mo agoHugging Face20LamTNguyen /ridgelora-cross-sensor-sd302d-f-to-m-20260825 RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment This public archive contains the leakage-controlled direct cross-sensor experiment used to evaluate whether Stage-2 synthetic target-sensor images help recognition on a physically different real sensor. Locked protocol Source/condition sensor: NIST SD302A device F. Target sensor: NIST SD302D device M. Identity: subject:finger-position; the same fingers exist across both collections. Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.imageimage-to-image1K<n<10K0 likes280 downloads1mo agoHugging Face21aeyxen /stem-diagrams STEM Diagrams 30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures) extracted from arXiv papers across six engineering fields, each with a source attribution and a quality score. Built by an LLM-curated pipeline and used to show that a small frozen-feature classifier can replace the paid LLM labeling gate. Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026) Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.imageimage-classification10K<n<100K0 likes250 downloads2mo agoHugging Face22EthnicErotic /phenotype-catalog Ethnic Erotic Phenotype Catalog A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations. Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research. What's in v6 Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.imagetext-classification10K<n<100K1 likes211 downloads21h agoHugging Face23marcelohaps /ijb-a IJB-A HF-ready This repo packages the IARPA Janus Benchmark-A (IJB-A) face recognition dataset in its CleanData layout, plus the official 10-split 1:1 verification and 1:N identification protocols. Unlike LFW-style benchmarks, IJB-A is template-based: each subject is represented by a template aggregating multiple still images and/or video frames. Protocol CSVs map every (template, file) row to a face annotation with bounding box, landmarks, and demographic attributes.… See the full description on the dataset page: https://huggingface.co/datasets/marcelohaps/ijb-a.imageimage-classification100K<n<1M0 likes209 downloads5mo agoHugging Face24Kanhaiyya /Sap_Kush_Med_Deepfakegated Sap_Kush_Med_Deepfake Dataset Paired medical-image forgery lineages across six modalities. Every lineage is one source image, one mask, one seed: the arms differ only in what was done inside the mask, so a comparison between arms isolates the manipulation rather than an encoding artefact. 3956 lineages, 30385 files, 8.49 GiB. v2 adds a removal arm grounded in human annotation for three more modalities (endoscopy, ultrasound, MRI). v1 had removal for CT only. What the… See the full description on the dataset page: https://huggingface.co/datasets/Kanhaiyya/Sap_Kush_Med_Deepfake.imageimage-classification1K<n<10K0 likes189 downloads12h agoHugging Face25Nima0Kamali /humancentric-scenes-ai HumanCentric-Scenes-AI A multimodal benchmark of 296 AI-generated human-centric scenes across four domains: CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise. Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/Nima0Kamali/humancentric-scenes-ai.imageimage-classificationn<1K2 likes185 downloads11d agoHugging Face26atharvadagaonkar /ImageNet-CJ JPEG Re-encoding Confound Control Dataset A controlled-experiment dataset that isolates one acknowledged-but-unmeasured confound in ImageNet-C. Hendrycks & Dietterich (Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, ICLR 2019, arXiv:1903.12261) save every corrupted image as a lightly compressed JPEG. The benchmark therefore never measures a corruption c applied to an image x in isolation — it measures JPEG(c(x)). This dataset lets you quantify how… See the full description on the dataset page: https://huggingface.co/datasets/atharvadagaonkar/ImageNet-CJ.imageimage-classification10K<n<100K0 likes181 downloads4mo agoHugging Face27star092304 /typhoon-intensity-classification Typhoon - Image Classification Dataset This dataset comes from PTIT AI Challenge and is organized for a multi-class image classification task focusing on tropical cyclone (typhoon) intensity estimation. Dataset Structure The directory structure is organized as follows: train/ ├── images/ │ ├── image1.jpg │ └── ... └── annotations.csv (only present in the train folder) The public_test and private_test sets are used to evaluate and score the… See the full description on the dataset page: https://huggingface.co/datasets/star092304/typhoon-intensity-classification.imageimage-classification1K<n<10K1 likes168 downloads4mo agoHugging Face28Nininkkka /Synth-Text-Eng-512x128 Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise. All samples are accompanied by a structured metadata.csv and a ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/Nininkkka/Synth-Text-Eng-512x128.imageimage-to-text10K<n<100K1 likes156 downloads12d agoHugging Face29PastaEvangelists /UK_Traffic_Sign_Inspection_Datasetimageimage-classification10K<n<100K0 likes154 downloads6mo agoHugging Face30yuyingzzz /jingchu-cultural-heritage 荆楚文化文物语义分析样本 这是用于审核字段设计和语义抽取质量的样本版本,共 116 条记录、35 个核心字段。 数据集另含 enrichment_trial_5 配置:从主表选取 5 条文物进行检索、图像观察与语义补充,共 40 个字段。主表内容未被覆盖。 湖北省博物馆:83 条 荆州博物馆:33 条 图片位于第 2 字段 image_url 两个古籍书影汇总页已拆分为 18 条单书记录 删除了当前来源完全无法填充的 creation_place 和 collection_number 删除派生检索字段 keywords;删除与 archaeological_site 高度重复的 provenance 原文与语义归纳分离;缺少来源的信息保持为空 所有记录目前均为 待人工复核,尚不是最终 604 条全量版本 数据文件 data/artifacts.csv:Dataset Viewer 使用的 116 条主表… See the full description on the dataset page: https://huggingface.co/datasets/yuyingzzz/jingchu-cultural-heritage.imageimage-classificationn<1K1 likes145 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.