datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cats-vs-dogs-sample
Dataset Card for Dataset Name
Subset of https://huggingface.co/datasets/microsoft/cats_vs_dogs, converted into FiftyOne dataset format.
This is a FiftyOne dataset with 5000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/cats-vs-dogs-sample.hard-intersection-multimodal-sample
Hard Intersection Multimodal Samples
Release Notes
Release
Description
v1.0.0
Initial public release.
v1.1.0
Added Unreal Engine assets.Fixed issues in the OpenDRIVE map data.Updated the README to improve documentation and usability.
v1.2.0
Added OpenCRG road surface elevation data.Updated 3DGS reconstruction assets.Updated the README to document OpenCRG support.
Dataset Summary
Hard Intersection Multimodal Samples is a curated… See the full description on the dataset page: https://huggingface.co/datasets/dynamic-maps/hard-intersection-multimodal-sample.STRI-Samples
Dataset Card for Smithsonian Tropical Research Institute (STRI) Samples
Dataset Summary
Dorsal images of butterfly wings collected by Owen McMillan and members of his lab at the Smithsonian Tropical Research Institute.
Full dataset will be 24,119 RGB images: Dorsal and Ventral images of separated wings. This sample contains 207 dorsal butterfly images used as part of the training data for Imageomics/butterfly_detection_yolo.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/STRI-Samples.farmerchat-image-samples
FarmerChat Crop Image Samples
A sample of farmer-submitted photographs from FarmerChat, an agricultural advisory service
used by smallholder farmers in India, Ethiopia, Kenya and Nigeria. Published by
Digital Green.
This release contains 6,089 records (5,957 distinct photographs; some
photographs belong to more than one category, see below) drawn from 7 categories
representing different outcomes of an automated crop diagnosis pipeline, sampled
across country, month, crop and… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/farmerchat-image-samples.dental-implant-surgery-sample
Dental Implant Surgery — Multimodal Annotated Video (Full-Mouth Three-Camera Sample)
A public preview of one complete full-mouth implant rehabilitation — both jaws in a
single session, six implants in the maxilla and six in the mandible with multi-unit
abutments — recorded by three synchronised cameras in a working operating room, with
the surgeons' own words aligned to the picture and every keyframe annotated across six
layers (L0–L5).
This repository is a showcase slice of… See the full description on the dataset page: https://huggingface.co/datasets/OralSurgery/dental-implant-surgery-sample.DataSeeds.AI-Sample-Dataset-DSD
DataSeeds.AI Sample Dataset (DSD)
This is a sample. For larger collections, custom production or annotation, write to sales@dataseeds.ai.
Dataset Summary
The DataSeeds.AI Sample Dataset (DSD) is a high-fidelity, human-curated computer vision-ready dataset comprised of 7,772 peer-ranked, fully annotated photographic images, 350,000+ words of descriptive text, and comprehensive metadata. While the DSD is being released under an open source license, a sister dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dataseeds/DataSeeds.AI-Sample-Dataset-DSD.game-assets-samplers
CyberMax game asset samplers (free)
Use this dataset
from datasets import load_dataset
ds = load_dataset("CyberMax-tools/game-assets-samplers", split="train")
df = ds.to_pandas()
print(df.shape)
df.head()
Want the full packs? Every sampler below is a small slice of a paid CyberMax pack with a live browser demo.
All 8 packs together: CyberMax Game Assets Mega Bundle — $37 · hub: https://cybermax.cybermax-tools.workers.dev/game-assets/?s=hf-ga
Original game… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/game-assets-samplers.temporal-aerial-cityline-construction-sample
CityLine — Temporal Aerial Construction Dataset (Sample)
Temporal Aerial Vision · Construction Progress · Multiview Geometry · San Jose, CA
CityLine is a multi-year aerial imagery sequence captured from a helicopter during the construction of a major mixed-use development in San Jose, California.This sample highlights multiple construction phases over time, with several oblique views per capture date.
The full (commercial) dataset contains hundreds of high-resolution images with… See the full description on the dataset page: https://huggingface.co/datasets/SharpShots/temporal-aerial-cityline-construction-sample.synthetic-australian-medical-documents-sample
Synthetic Australian Medical Documents - Sample
A 50-document free sample of a 5,000-document library of synthetic Australian medical PDFs. PHI-free. Modelled on Australian healthcare documentation. Pre-labelled with structured ground truth and pixel-precise bounding boxes. Released under CC-BY-NC 4.0 for evaluation and non-commercial research.
See Pricing & licensing below.
What's in this sample
Field
Value
Documents
50
Document types
29 (of 45 in full… See the full description on the dataset page: https://huggingface.co/datasets/RootCauseAnalytics/synthetic-australian-medical-documents-sample.scallop_mosaic_640_quantization_sample
Scallop YOLOv5s Mosaic 640 - Quantization Sample
A classless 1000-train / 1000-val image subset of the tiled 3x3 640px mosaic dataset designed specifically for RKNN/tflite/ONNX representative quantization calibration on edge devices (like the RV1126 Aura).
Attribution & License
This dataset is a derivative work based on the University of St Andrews King Scallop dataset.
Original DOI: 10.5281/zenodo.10156830
In accordance with the original dataset's terms, this… See the full description on the dataset page: https://huggingface.co/datasets/FishingROV/scallop_mosaic_640_quantization_sample.uk-property-inspection-sample
miProgram UK Residential Property Inspection Sample
A 423-image labelled sample of professional UK residential property inspection photography, drawn from real production inspections carried out through the miProgram platform, prepared for evaluation by AI training-data buyers and data-licensing partners.
This is a sample. Full commercial dataset (approximately 120 million professionally-labelled images and growing) available under separate commercial licence — contact… See the full description on the dataset page: https://huggingface.co/datasets/miprogram/uk-property-inspection-sample.M-Attack-V2-Adversarial-Samples
M-Attack-V2 Adversarial Samples
Adversarial image samples generated by M-Attack-V2, from the paper:
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
arXiv:2602.17645 | Project Page | Code
Dataset Structure
├── epsilon_8/ # 100 adversarial images (ε = 8/255)
│ ├── 0.png
│ ├── 1.png
│ ├── ...
│ └── metadata.csv
└── epsilon_16/ # 100 adversarial images (ε = 16/255)
├── 0.png
├── 1.png
├── ...
└──… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-LLM/M-Attack-V2-Adversarial-Samples.bakkhali-estuary-high-tide-sample
Bakkhali River Estuary — High Tide Boat Survey (Free Sample)
50 GPS-tagged coastal images from a single high-tide boat survey of the Bakkhali River
estuary, Khurushkul, Cox's Bazar, Bangladesh.
By Golam Rob — www.golamrob.com
✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)".
Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution.
📸 These 50 images are a small taste of a 200,000–300,000 image personal library of… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/bakkhali-estuary-high-tide-sample.WildFake-Sample
WildFake-Sample
A 30,000-image sample of WildFake (Hao et al., AAAI 2025,
arXiv:2402.11843;
original dataset),
covering generators and real-image sources outside DDA/SID — a held-out
generalization slice, not a copy of the full ~3.6M-image dataset. All credit
for the images goes to WildFake's original authors. Built for
Buxt-Codes/AIGI-Detection
(branch LoRC-PC) — see that repo's HANDOFF.md for the evaluation
methodology and results.
Composition
Fake (19,500):… See the full description on the dataset page: https://huggingface.co/datasets/buxtcodes/WildFake-Sample.khurushkul-pond-water-lily-sample
Pink Water Lily & Water Hyacinth — Khurushkul Pond, Bangladesh
100 GPS-tagged freshwater wetland images from a single pond survey in Khurushkul, Cox's Bazar, Bangladesh. By Golam Rob — www.golamrob.com
✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)". Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution.
📸 These 100 images are a small taste of a 200,000+ image personal library of coastal, tidal, and freshwater… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/khurushkul-pond-water-lily-sample.honeybee-samples
HoneyBee Sample Files
Sample data and resource files for the HoneyBee framework — a scalable, modular toolkit for multimodal AI in oncology.
These files are used by the HoneyBee example notebooks (clinical, pathology, radiology) and by HoneyBee's molecular processing code at runtime (Hugo_symbols.tsv is fetched on first use of DNA mutation preprocessing).
Paper: HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models… See the full description on the dataset page: https://huggingface.co/datasets/Lab-Rasool/honeybee-samples.LMOD-Plus-Sample
LMOD+ — Sample Subset
A 1,076-instance sample of LMOD+, a large-scale multimodal
ophthalmology benchmark for developing and evaluating multimodal large language models (MLLMs).
This repository is a preview subset intended for quickly inspecting the data format, prototyping
evaluation harnesses, and running smoke tests. The full benchmark contains 32,633 instances across
12 ophthalmic conditions and 5 imaging modalities.
📄 Paper: ACM Transactions on Computing for Healthcare… See the full description on the dataset page: https://huggingface.co/datasets/Euanyu/LMOD-Plus-Sample.PRISM-Dataset-Sample
PRISM Sample: Polarimetric Road-surface Intelligent Sensing and Measurement Dataset
Anonymous submission to NeurIPS 2026 Evaluations & Datasets Track.
This is a representative sample of the PRISM dataset, designed to enable reviewers and researchers to inspect data quality without downloading the full ~1.6 TB dataset.
Why a sample dataset?
The full PRISM dataset contains 47,098 time-synchronized frames across 41 sessions. This sample provides:
Quick quality inspection:… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-2026-PRISM/PRISM-Dataset-Sample.SAGE-sample
SAGE — Sample (10 crops preview)
A focused 10-crop sample of the full tirtho149/SAGE dataset. Each (crop, disease) class contributes up to 10 images (fixed seed for reproducibility).
For the full ~280 GB dataset, see tirtho149/SAGE.
Crops included
Crop
Disease classes
Images
Apple
36
209
Corn
86
679
Cotton
14
74
Mango
9
55
Potato
37
306
Rice
39
261
Soybean
58
498
Sugarcane
17
170
Tomato
48
365
Wheat
53
374
TOTAL
1,848… See the full description on the dataset page: https://huggingface.co/datasets/tirtho149/SAGE-sample.aigc-sample
AIGC Dataset — NovelAI Generations
⚠️ Contains NSFW content. Images and prompts have not been filtered. Many are sexually explicit.
1,562,685 AI-generated images (≈332 GB) made with NovelAI's image models (V3 → V4 → V4.5 → V5)
between 2023-11 and 2026-09 through a private Discord bot (Kohaku-NAI).
Every image has its full generation metadata (model, sampler, steps, CFG, seed, prompt, …) and an exact
generation timestamp. That makes the set usable for:
AI-generated image… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/aigc-sample.hlt006-sample
HLT-006 — Synthetic Medical Imaging Dataset (Sample Preview)
A free, schema-identical preview of the full HLT-006 commercial product from XpertSystems.ai.
A fully synthetic medical imaging dataset combining study-level metadata, COCO-format bounding box and segmentation annotations, DICOM tag fields, and structured radiologist reports. Calibrated to NIH ChestX-ray14, LIDC-IDRI, BraTS, MRNet, and ACR RADS standards across CXR (Chest X-ray), CT (Chest/Abdomen/Head), and MRI… See the full description on the dataset page: https://huggingface.co/datasets/xpertsystems/hlt006-sample.samuel-and-audrey-photography-metadata-archive
Samuel & Audrey Photography Metadata Archive
This dataset contains a structured metadata archive for the Samuel & Audrey Media Network travel photography collection hosted on SmugMug.
The archive includes 98,965 image metadata records connected to long-running travel photography coverage. Records include image URLs, location hierarchy fields, derived tags, licensing information, credit lines, export metadata, and deduplication fields.
This dataset provides metadata and source URLs… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-photography-metadata-archive.bharatanatyam-mudra-dataset
Bharatanatyam Mudra Dataset
Dataset Description
The Bharatanatyam Mudra Dataset contains 28,431 images of hand gestures (mudras) from Bharatanatyam, a classical Indian dance form. The dataset was collected from 15 volunteers in a studio environment and includes both single-hand and double-hand gestures.
Dataset Statistics
Total Images: 28,431
Single Hand Gestures (Asamyukta Hastas): 15,396 images across 29 classes
Double Hand Gestures (Samyukta Hastas): 13,035… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/bharatanatyam-mudra-dataset.mfg010-sample
MFG-010 — Manufacturing Defects Dataset (Sample)
A schema-identical preview of MFG-010, the XpertSystems.ai synthetic
defect events with visual-inspection ML metadata dataset for AOI
(Automated Optical Inspection) ML training, FMEA RPN modeling,
Ishikawa root cause classification, CAPA workflow simulation, and
defect-cohort quality engineering research. The full product covers
10,000-100,000 records. This sample is HF-sized at 3,000 records.
Built by XpertSystems.ai — Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/xpertsystems/mfg010-sample.pd-defect
Author-Maintained Fork
This repository is an author-maintained fork of the original PD-Defect dataset repository hosted under the EPDCL organization. The original repository is available at EPDCL/pd-defect.
This repository is maintained by the dataset authors to provide continued public access to the dataset, its documentation, and its version history. The dataset and associated materials remain subject to the license and attribution requirements specified in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/sampath-balaji/pd-defect.central-europe-rural-landscape-dataset-sample
SAMPLE VERSION (Preview Subset)
This repository contains the preview sample (45 images) of the Central European Rural Landscape Dataset (v1.0).
The full dataset (945 images, class-wise ZIP archives, full annotation set) is available separately under a commercial license.
Full version repository:
👉 https://huggingface.co/datasets/batris-data/central-europe-rural-landscape-dataset-full
For licensing inquiries:
batris.sro@gmail.com
Central European Rural Landscape Dataset… See the full description on the dataset page: https://huggingface.co/datasets/batris-data/central-europe-rural-landscape-dataset-sample.jump-sample
cp-bg-bench preview — jump
Compact preview of the jump dataset from the cp-bg-bench
benchmark. Stratified subset of the full release; designed so that the
held-out-batch perturbation-recall and cp_measure-prediction evals can
be reproduced end-to-end against this small slice alone.
Cells
913
Wells
69
Perturbations
33
Held-out batch
source_4
Views
crops, crops_density, seg, seg_density
What's in this repo
crops/ # HF dataset… See the full description on the dataset page: https://huggingface.co/datasets/cp-bg-bench-anon/jump-sample.MV-VDB-photos-small
MV-VDB-photos-small
Media Vault - Vector Database Photos (Small)
A curated collection of 11,000 images from various computer vision datasets, designed for testing internal mechanisms in the Media Vault Vector Database system. This is the first small-scale dataset (targeting 10K samples, with NSFW split totaling 11K) for validation and testing purposes.
Dataset Structure
The dataset contains two splits:
sfw: All non-NSFW images (~10,000 images)
x_nsfw: Only NSFW images… See the full description on the dataset page: https://huggingface.co/datasets/SamoXXX/MV-VDB-photos-small.core_sample_image_data
🖼 Soil Core Sample Image Data
This dataset contains 3604 labeled images of soil core samples for image classification.
📌 Dataset Summary
Images: Squared images (300x300 pixels), cropped from full-scale high-resolution images of soil core samples.
Labels: hb, nb (terms according to DIN 4023:2006-02).
Format: Hugging Face datasets.Dataset with Image() feature.
Split: Train / val / test split is performed with a ratio of 0.8 / 0.1 / 0.1, whereas the samples are stratified… See the full description on the dataset page: https://huggingface.co/datasets/grano1/core_sample_image_data.labelled-samples
WHA Spell Simulator Glyphs
Crowdsourced handwriting samples of signs and sigils from the fan-made
Witch Hat Atelier spell simulator. Contributors drew each
glyph freehand in the project's Sample Maker tool; every sample was
human-reviewed and only approved samples are included. Strokes are
simplified (polling-rate invariant) and normalised to the 0..1 range,
preserving aspect ratio.
vector config
One record per sample:
id — content hash of the raw sample… See the full description on the dataset page: https://huggingface.co/datasets/wha-spell-simulator/labelled-samples.
