Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes31k downloads4mo agoHugging Face02hanchong /synthetic-infrared-maritime-vessel-dataset-flux2klein Synthetic Infrared Maritime Vessel Dataset (FLUX2-Klein) RGB maritime vessel images translated to synthetic infrared via a DreamBooth-finetuned FLUX.2-Klein-4B. Layout synthetic-infrared-maritime-vessel-dataset-flux2klein/ ├── in-distribution/ │ ├── train/{C00,C02,...}/*.jpg │ ├── val/{C00,C02,...}/*.jpg │ ├── test/{C00,C02,...}/*.jpg │ ├── labels.txt │ └── selected-metadata-{train,val,test}.json └── out-of-distribution/ ├── val/{C01,C03,C06… See the full description on the dataset page: https://huggingface.co/datasets/hanchong/synthetic-infrared-maritime-vessel-dataset-flux2klein.imageimage-classification100K<n<1M0 likes25k downloads2mo agoHugging Face03zalando-datasets /fashion_mnist Dataset Card for FashionMNIST Dataset Summary Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. We intend Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for benchmarking machine learning algorithms. It shares the same image size and structure of training and testing… See the full description on the dataset page: https://huggingface.co/datasets/zalando-datasets/fashion_mnist.imageimage-classification10K<n<100K67 likes21k downloads2y agoHugging Face04alkzar90 /NIH-Chest-X-ray-datasetThe NIH Chest X-ray dataset consists of 100,000 de-identified images of chest x-rays. The images are in PNG format. The data is provided by the NIH Clinical Center and is available through the NIH download site: https://nihcc.app.box.com/v/ChestXray-NIHCCimage-classification100K<n<1M63 likes15k downloads2y agoHugging Face05perturb-ai /efficientnet-v2-l-adv-dataset Perturb Adversarial Images Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the Perturb network. Each row is one clean image together with all of its verified adversarial versions: images that are imperceptibly different from the original (L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction. This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.imageimage-classification10K<n<100K0 likes7.1k downloads13m agoHugging Face06Voxel51 /MPII_Human_Pose_Dataset Dataset Card for MPII Human Pose MPII Human Pose dataset is a state of the art benchmark for evaluation of articulated human pose estimation. The dataset includes around 25K images containing over 40K people with annotated body joints. The images were systematically collected using an established taxonomy of every day human activities. Overall the dataset covers 410 human activities and each image is provided with an activity label. Each image was extracted from a YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/MPII_Human_Pose_Dataset.imageimage-classification10K<n<100K17 likes5.5k downloads2y agoHugging Face07hanchong /real-infrared-maritime-vessel-dataset Real Infrared Maritime Vessel Dataset Real infrared imagery of maritime vessels. The dataset is provided in three forms — full-frame detection images, per-object classification crops, and a hand-curated subset. Classes (7): liner, bulk carrier, warship, sailboat, canoe, container ship, fishing boat. Layout real-infrared-maritime-vessel-dataset/ ├── original/ Full-frame IR images + XML bounding-box labels (detection) │ ├── images/{train,test}/*.jpg… See the full description on the dataset page: https://huggingface.co/datasets/hanchong/real-infrared-maritime-vessel-dataset.imageimage-classification10K<n<100K2 likes3.5k downloads2mo agoHugging Face08a2015003713 /military-aircraft-detection-dataset Military Aircraft Detection Dataset Military aircraft detection dataset in COCO and YOLO format. The dataset was initially developed exclusively for military aircraft detection, but was later expanded to include commercial airliners for a broader and more challenging detection task. The dataset contains 103 military aircraft types and 11 commercial airliner types. Military aircraft: A10, A400M, AG600, AH64, AKINCI, AV8B, An124, An22, An225, An72, B1, B2, B21, B52, Be200, C1… See the full description on the dataset page: https://huggingface.co/datasets/a2015003713/military-aircraft-detection-dataset.imageobject-detection10K<n<100K3 likes3.4k downloads17d agoHugging Face09rwcuffney /autotrain-data-pick_a_card AutoTrain Dataset for project: pick_a_card Dataset Description This dataset has been automatically processed by AutoTrain for project pick_a_card. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "image": "<224x224 RGB PIL image>", "target": 0 }, { "image": "<224x224 RGB PIL image>", "target": 0 }] Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/rwcuffney/autotrain-data-pick_a_card.image-classification1 likes3.2k downloads4y agoHugging Face10zr-zhang /MLLM-Generated-Image-Detection-Dataset MLLM-Generated Image Dataset This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection. Paper | Code Dataset Summary We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.imageimage-classification10K<n<100K1 likes3.1k downloads4d agoHugging Face11Voxel51 /Describable-Textures-Dataset Dataset Card for Describable Textures Dataset This is a FiftyOne dataset with 5640 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/Describable-Textures-Dataset") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Describable-Textures-Dataset.imageimage-classification1K<n<10K4 likes2.7k downloads2y agoHugging Face12LSL-datasets /TALKtoME TALKtoME: Educational Materials for Speech and Language Acquisition in Autism This dataset contains image and video samples of action verbs and verb+noun pairs. It is designed to support machine learning tasks related to visual understanding, action recognition, and language grounding. Dataset Structure The dataset contains the following main folders: action_verbs/images/: image samples organized by action verb categories. action_verbs/videos/: video samples organized by… See the full description on the dataset page: https://huggingface.co/datasets/LSL-datasets/TALKtoME.imageimage-classification1K<n<10K1 likes2.2k downloads5mo agoHugging Face13Kaphathy /Dataset MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records 1. Executive Summary & Repository Overview The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.textimage-classificationn<1K2 likes2.1k downloads12d agoHugging Face14Rajarshi-Roy-research /Defactify_Image_Dataset Defactify_Image_Dataset This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection. 📝 Dataset Description Dataset Summary The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.imageimage-classification10K<n<100K22 likes1.9k downloads5mo agoHugging Face15claytonsds /UTM_Dataset Dataset Description This dataset was created to support Article: VisionGauge: a computer vision model to detect and read U-tube manometers and is available on GitHub It consists of images of U-tube manometers constructed using a transparent PVC water level hose (5/16" × 1 mm) and flexible measuring tapes of different colors, each with a length of 150 cm (60 inches). The manometric fluids represented in the dataset include water, oil, and dyed water. The dataset is intended… See the full description on the dataset page: https://huggingface.co/datasets/claytonsds/UTM_Dataset.imageimage-classification10K<n<100K0 likes1.8k downloads4d agoHugging Face16ioandanielc /sph_dataset SPH-Simulated LPBF Melt-Pool Dataset Single-track laser powder bed fusion (LPBF) melt-pool simulations for Ti-6Al-4V, produced with the LAMAS smoothed-particle-hydrodynamics solver. 241 simulations sampled uniformly i.i.d. over a 4D process-parameter cube (laser power, scan speed, laser spot radius, substrate temperature), spanning conduction, transition, and keyhole regimes. Companion to the NeurIPS 2026 Evaluations & Datasets Track submission A Simulation-Based Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/ioandanielc/sph_dataset.image-classification100K<n<1M0 likes1.8k downloads3mo agoHugging Face17ctmedtech /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/DDR-dataset.imageimage-segmentation10K<n<100K2 likes1.8k downloads11mo agoHugging Face18BJyotibrat /Masoud-Nickparvar-Brain-Tumor-MRI-Dataset Dataset Card for Brain Tumor MRI Dataset A dataset of 7,200 human brain MRI images, labeled into four classes — glioma, meningioma, pituitary tumor, and no tumor — for training and evaluating brain tumor classification models. This is a direct upload of the Brain Tumor MRI Dataset originally published on Kaggle by Masoud Nickparvar. Dataset Details Dataset Description This dataset combines MRI images from three source datasets — figshare, the… See the full description on the dataset page: https://huggingface.co/datasets/BJyotibrat/Masoud-Nickparvar-Brain-Tumor-MRI-Dataset.imageimage-classification1K<n<10K1 likes1.7k downloads2mo agoHugging Face19Xanadu00 /autotrain-data-galaxy_classification AutoTrain Dataset for project: galaxy_classification Dataset Description This dataset has been automatically processed by AutoTrain for project galaxy_classification. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "image": "<256x256 RGB PIL image>", "target": 0 }, { "image": "<256x256 RGB PIL image>", "target": 0 }]… See the full description on the dataset page: https://huggingface.co/datasets/Xanadu00/autotrain-data-galaxy_classification.imageimage-classification1 likes1.7k downloads3y agoHugging Face20Voxel51 /qualcomm-interactive-video-dataset Dataset Card for Qualcomm Interactive Video Dataset This is a FiftyOne dataset with 2900 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/qualcomm-interactive-video-dataset") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/qualcomm-interactive-video-dataset.videoimage-classification1K<n<10K2 likes1.6k downloads10mo agoHugging Face21vivekvar /cctv-datasets CCTV Datasets for helmet detection + ANPR Training and evaluation data used by vivekvar/helmet-v5 and vivekvar/helmet-v4. Source: Andhra Pradesh RTGS CCTV feeds (public road cameras). All crops and frames are from motorcycle traffic scenes. Folders Folder Contents Purpose merged_v3/ YOLO-format dataset (data.yaml + train/valid/test) Bike + rider detection training clean_merged_data/ Cleaned / deduped crop set Base training data for v4 extra_khadatkar/… See the full description on the dataset page: https://huggingface.co/datasets/vivekvar/cctv-datasets.imageobject-detection10K<n<100K0 likes1.6k downloads6mo agoHugging Face22Voxel51 /scanned-images-dataset-for-ocr-and-vlm-finetuning Dataset Card for scanned_images_dataset This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.imageimage-classification1K<n<10K2 likes1.6k downloads8mo agoHugging Face23Naiscorp /car-damage-dataset Car Damage Images A raw image collection for vehicle damage assessment. Unlabeled: these images have no annotations yet and are intended as source material for labeling or pre-training. Structure images/ 001/ images_001.jpg images_002.jpg ... 002/ ... ... 683 folders thumbnails/ 001/ thumbnail_001.jpg thumbnail_002.jpg ... ... 683 folders Branch Files Folders Size… See the full description on the dataset page: https://huggingface.co/datasets/Naiscorp/car-damage-dataset.imageimage-classification10K<n<100K4 likes1.5k downloads29d agoHugging Face24ISU-Test /isu-challenge-dataset Dataset Card for ISU Challenge Dataset Dataset Summary ISU Challenge Dataset is a multi-modal in-cabin automotive dataset with a controlled synthetic core and paired real-reference scenes. The synthetic core contains 1,000 synchronized Blender-rendered samples with: RGB render depth (EXR and PNG) instance segmentation canny edge map structured scenario labels An additional 60 paired real-reference scenes occupy sample_00000 through sample_00059. Each paired… See the full description on the dataset page: https://huggingface.co/datasets/ISU-Test/isu-challenge-dataset.imageimage-segmentation1K<n<10K2 likes1.5k downloads16d agoHugging Face25MelanieCo /Capillary-Dataset Capillary dataset Paper: Capillary Dataset: A dataset of nail-fold capillaries captured by microscopy for diabetes detection Github: https://github.com/urgonguyen/Capillarydataset.git The dataset are structured as follows: Capillary dataset ├── Classification ├── data_1x1_224 ├── data_concat_1x9_224 ├── data_concat_2x2_224 ├── data_concat_3x3_224 ├── data_concat_4x1_224 └── data_concat_4x4_224 ├── Morphology_detection… See the full description on the dataset page: https://huggingface.co/datasets/MelanieCo/Capillary-Dataset.imageimage-classification10K<n<100K0 likes1.5k downloads3mo agoHugging Face26sdu-chenxin /MDS-Bench-data MDS-Bench Raw Data Raw input data for MDS-Bench, a benchmark that tests whether AI agents can organize heterogeneous medical imaging datasets into a unified image + JSON format. It contains 100 public medical imaging datasets (CT, MRI, microscopy, X-ray, ultrasound, endoscopy, ...; DICOM, NIfTI, TIFF, H5, MAT, ...), about 2 TB in total. Layout origin_dataset/<dataset_name>/data/... # 8 datasets second_dataset/<dataset_name>/data/... # 34 datasets… See the full description on the dataset page: https://huggingface.co/datasets/sdu-chenxin/MDS-Bench-data.image-segmentation0 likes1.5k downloads8d agoHugging Face27MITLL /LADI-v2-dataset Dataset Card for LADI-v2-dataset Dataset Summary : v2 The LADI-v2 dataset is a set of aerial disaster images captured and labeled by the Civil Air Patrol (CAP). The images are geotagged (in their EXIF metadata). Each image has been labeled in triplicate by CAP volunteers trained in the FEMA damage assessment process for multi-label classification; where volunteers disagreed about the presence of a class, a majority vote was taken. The classes are: bridges_any… See the full description on the dataset page: https://huggingface.co/datasets/MITLL/LADI-v2-dataset.imageimage-classification1K<n<10K8 likes1.3k downloads2y agoHugging Face28cs5242-hateful-memes /hateful-memes-data Hateful Memes (CS5242 submission mirror) Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020) used for reproducibility of our CS5242 (NUS) submission. Contents img/ — 10,000 PNG images of memes train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540), test_seen.jsonl (1,000), test_unseen.jsonl (2,000) Provenance This mirror merges two existing mirrors of the original Meta release: Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.imageimage-classification10K<n<100K2 likes1.2k downloads6mo agoHugging Face29Abhisheksvnit /ElectraAI-Dataset-v4 ElectraAI Dataset v4 1M+ record ECE/VLSI/Analog/Embedded dataset for LLM and vision model fine-tuning. Statistics Stream Records Description A: NGSpice Netlists 100,000 SPICE netlists + analyses B: Labeled Circuit Images 100,000 PNG + JSON (schemdraw) C: QA Pairs (ShareGPT) 700,000 Multi-turn ECE instruction D: YOLO Annotations 100,000 PNG + YOLO TXT (20 classes) Total 1,000,000 Repo layout (sharded directories) Images and… See the full description on the dataset page: https://huggingface.co/datasets/Abhisheksvnit/ElectraAI-Dataset-v4.text-generation1M<n<10M0 likes1.2k downloads3mo agoHugging Face30anonymousllbench /llbench-dataset LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute. LL-Bench is a large-scale, human-preference benchmark for evaluating low-level vision restoration in the era of large generative models (LGMs). It compares 10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.imageimage-to-image100K<n<1M0 likes1.2k downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.