datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clamp-benchmark
CLAMP: A Sim-to-Real Benchmark for Closed-Loop Kinematic Pose Estimation and Assembly Reasoning
Closed-Loop Assembly and Mechanism Perception
Accepted at NeurIPS 2026 (Evaluations & Datasets Track)
Kevin Murray1, Randolph Beauregard Robert III2, Petar Z Duric1, Zoran Duric3
1Overlab LLC 2AVA Labs 3George Mason University
📦 Code: https://github.com/overlab-kevin/clamp
Your browser does not support the video tag.
Overview of all 210 labeled real test scenes.… See the full description on the dataset page: https://huggingface.co/datasets/clamp-benchmark/clamp-benchmark.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.gently-perception-benchmark
Gently Perception Agent Benchmark
Light-sheet microscopy volumes of C. elegans embryo development, intended
for evaluating vision-based perception agents on embryo stage classification.
The dataset has two tiers:
Annotated benchmark set (embryo_1–embryo_8) — human ground-truth
stage transitions. Use this for evaluation.
Unannotated corpus (embryo_9–embryo_105) — 97 additional real embryo
timelapses with no human labels, provided for developing and stress-
testing perception… See the full description on the dataset page: https://huggingface.co/datasets/gently-project/gently-perception-benchmark.ImageTime_Benchmark
ImagineTime Benchmark
This dataset repository contains the public benchmark assets for ImagineTime, released with the paper “Can Image Models Imagine Time?”
Paper: arXiv:2606.10620
ImagineTime evaluates whether image generation models can produce ordered 2x2 motion sheets with coherent entities, spatial relations, state transitions, interactions, and task constraints.
Contents
cases/
750 benchmark cases. Each case includes process specs, prompts… See the full description on the dataset page: https://huggingface.co/datasets/Xin-Rui/ImageTime_Benchmark.geo-benchmarks
im2gps3k and yfcc4k
Two standard image-geolocation test sets, packaged as zips.
File
Contents
im2gps3k.zip
im2gps3k/im2gps3k_places365.csv (columns IMG_ID, AUTHOR, LAT, LON, ...; 2,997 rows) and im2gps3k/images/ (3,000 photos)
yfcc4k.zip
yfcc4k/yfcc4k.csv (columns IMG_ID, OwnerNSID, LAT, LON, ...; 4,536 rows), yfcc4k/yfcc4k.txt (raw YFCC100M metadata) and yfcc4k/images/ (4,536 photos)
Unzip both into one directory to get the layout <root>/{im2gps3k,yfcc4k}/....… See the full description on the dataset page: https://huggingface.co/datasets/kinghorton42/geo-benchmarks.artist-style-benchmark
Artist Style Benchmark
A benchmark dataset of 34903 anime-style illustrations generated with
AnimaImagine, each using a
different Danbooru artist tag at fixed prompt/seed settings.
ASR Ranker Style Browser
This dataset also serves as the image host for
ASR Ranker, an interactive browser for comparing and
ranking artist styles.
Browser: https://ranker.kuronet.top/browser
Dataset: https://huggingface.co/datasets/Moeblack/artist-style-benchmark
选出最 hot 的画风吧。
总榜只有 trusted… See the full description on the dataset page: https://huggingface.co/datasets/Moeblack/artist-style-benchmark.PRISM_Benchmark
PRISM: PhotoRealistic Image Synthesis and Manipulation
Dataset repository for the paper "The PRISM benchmark: PhotoRealistic Image Synthesis and Manipulation to detect generated images" — Bartolucci, Salti, Lisanti.
Real images
The corresponding real images can be downloaded separately from COCO:
Split
Source
Test set
COCO 2017 Val
Training set
COCO 2017 Train
Citation
@article{BARTOLUCCI2026104826,
title = {The PRISM… See the full description on the dataset page: https://huggingface.co/datasets/oppiliF/PRISM_Benchmark.geofm-agriculture-benchmark
GeoFM Agriculture Benchmark
Sample data and fine-tuned weights accompanying our ACM SIGSPATIAL 2026 paper, released so
other researchers can run inference with SatMAE, Prithvi, and SpectralGPT on our
multi-temporal crop segmentation and change-detection tasks.
This is not the full training dataset — it's a set of representative chips per region/model
plus the fine-tuned checkpoints, enough to run and sanity-check inference end-to-end.
Contact the authors if you need the complete… See the full description on the dataset page: https://huggingface.co/datasets/sanmay4119/geofm-agriculture-benchmark.GeneLab_BPS_BenchmarkData
Dataset Card for Dataset GeneLab_BPS_BenchmarkData
Dataset Details
This dataset is a version of the Biological and Physical Sciences (BPS) Microscopy Benchmark Training Dataset managed by NASA and hosted on an S3 Bucket here: https://registry.opendata.aws/bps_microscopy/
Fluorescence microscopy images of individual nuclei from mouse fibroblast cells, irradiated with Fe particles or X-rays with fluorescent foci indicating 53BP1 positivity, a marker of DNA damage.… See the full description on the dataset page: https://huggingface.co/datasets/kenobi/GeneLab_BPS_BenchmarkData.CTTA-AD-Benchmarks
CTTA-AD Benchmarks
Dataset collection for CTTA-AD: Continual Test-Time Adaptation for Unified Few-Shot Visual Anomaly Detection (AAAI 2027 submission).
Datasets
Dataset
Domain
Categories
Train Normal
License
MVTec-AD
Industrial
15
209–391 per category
CC BY-NC-SA 4.0
VisA
Industrial
12
400–905 per category
CC BY-NC-SA 4.0
MVTec-LOCO
Logical
5
varies
CC BY-NC-SA 4.0
BrainMRI
Medical
1
7,500
Research only
LiverCT
Medical
1
1,542
Research only… See the full description on the dataset page: https://huggingface.co/datasets/Hammadhaideerr/CTTA-AD-Benchmarks.watermark-benchmark-v18
Multi-Class Watermark & Camera Stamp Dataset (Round 18)
This repository contains the complete dataset, augmentation assets, and real-world evaluation benchmarks used to train the Champion 3-Class Watermark Classifier (wm_3class_v18_scratch.pt).
The dataset addresses a critical problem in media ingestion: automatically excluding produced media, broadcast stills, and stock photography without falsely excluding authentic personal photographs (0.00% false alarms on personal photos… See the full description on the dataset page: https://huggingface.co/datasets/langminer/watermark-benchmark-v18.ppe-benchmark-eval
PPE Benchmark Eval Set (v1)
A held-out, human-verified benchmark for evaluating vision-language models on
personal protective equipment (PPE) detection — specifically hardhat and
safety-vest presence — framed as a VQA-style classification task.
What this is
96 images, balanced 24/24/24/24 across the four hardhat × vest combinations
(yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of
the karabuk-university PPE dataset
on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.ANT-A-Benchmark
📊 ANT-A Benchmark
This directory contains the ANT-A Benchmark in HuggingFace-compatible Parquet format, ready for upload to the HuggingFace Hub.
ANT-A is a benchmark of 7 real-world annotation tasks (10,500 samples, 5 annotators each), covering text and multimodal modalities. All 5 annotations are preserved per sample — no forced majority vote — supporting research on soft labels and annotator disagreement.
🗂️ Dataset Structure
Each task is provided as two… See the full description on the dataset page: https://huggingface.co/datasets/xiapk7/ANT-A-Benchmark.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,355,161 human responses, collected with the
Rapidata Python SDK, comparing how well 30 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.CoVAtt-Benchmark
CoVAtt-Benchmark
A large-scale benchmark for generated-image attribution: given an image
that was produced by some text-to-image model, decide which model
produced it, and decide whether it came from a model the system has never
seen before.
What this dataset is
This is the dataset used to train and evaluate CoVAtt
(Content-Based Verification for Attribution of AI-Generated Images,
BMVC 2026). CoVAtt is a Siamese network that takes a pair of images and
predicts… See the full description on the dataset page: https://huggingface.co/datasets/kenyag/CoVAtt-Benchmark.doc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.document-processing-benchmark
Document Processing Benchmark
8 public document datasets (receipts, invoices, forms, bank statements,
multi-page docs, contracts) normalized into one parquet schema. Each row
has the document, ground-truth annotations, and per-row token/latency/cost
numbers from real API calls to one or more reference models. You can
read off a target's cost/latency/quality without re-running it.
from datasets import load_dataset
ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.Reference-Update-Benchmark
COVER-Fish Reference-Update Benchmark
This benchmark studies how reference-corpus updates change a frozen recognition
system, instantiated on fine-grained fish identification. It publishes the
versioned control plane behind COVER-Fish: row-level manifests, gallery states,
taxonomy, tensor bindings, transition evidence, protocols and dependency locks.
The benchmark does not duplicate the 83 GB Full Payload Archive. Large source
archives and frozen tensors remain in the immutable… See the full description on the dataset page: https://huggingface.co/datasets/COVER-Fish/Reference-Update-Benchmark.scope-benchmark
SCOPE Benchmark
Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories.
The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE.
When you chain a language model and a vision model together, how do you know which one failed?
Contents
scope-benchmark/
scope_541.csv… See the full description on the dataset page: https://huggingface.co/datasets/HindsboNikolaj/scope-benchmark.poi-benchmark
POI Benchmark: Multi-City Multimodal Points of Interest
A large-scale multimodal benchmark pairing Points of Interest (POIs) with street-view imagery, aerial grid photos, and satellite imagery across 10 major cities on 3 continents.
Cities
Beijing, Chengdu, Guangzhou, Hong Kong, Shanghai, Shenzhen, London, Melbourne, New York, Sydney.
Contents
Path
Size
Type
Description
metadata_aligned.tar
8.6 GB
11 JSON files
Enriched & aligned POI metadata per… See the full description on the dataset page: https://huggingface.co/datasets/sukiewang/poi-benchmark.Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace.
The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels.
The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.Face_Generation_Benchmark
Rapidata Human Face Generation Alignment
This T2I dataset contains over ~22'000 human responses, collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating 12 different image generation models on which one can generate faces more accurately.
The question that the annotators get asked is: "Which Image follows the description of the human better?"
To evaluate your own models and create leaderboard check out our… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Face_Generation_Benchmark.document-classification-benchmark
Document Classification Benchmark (open-vocab, zero-shot)
Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot,
open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked
into a head. Test split only; not for training. Every image is drawn from a permissively-licensed,
redistributable source.
Powers the
document-classification-leaderboard
and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.medical-imaging-model-evaluation-benchmark
医学影像多模态模型评测集(精选示例版)
这是一个面向医学多模态大模型的高质量影像评测集,专门测试模型能否把“看见影像”进一步转化为可解释、可复核、符合临床语境的判断与表达。数据将医学影像与患者描述、病史摘要、检查信息或结构化临床资料配对,覆盖从影像分类、报告生成,到鉴别诊断、治疗方案和胸片质量控制的完整评测链路。
本次公开版本从 2026-07-22 质检通过产物中整理而来,按每个子集最多 50 题进行分层抽样;题量不足 50 的影像质量控制子集完整保留。因此,公开版本包含 5 个任务子集、213 题和 455 个配套影像文件,适合作为医学视觉语言模型的快速对比集、回归测试集和研究教学样例。
数据集亮点
多模态对齐:每条样例同时提供影像和结构化的 question、answer、explanation,支持检查视觉理解、临床语义整合与解释质量。
任务覆盖完整:从“影像是什么”到“如何描述、如何鉴别、如何处置”,并加入真实影像工作流中的胸片质量控制任务。
影像类型丰富:覆盖 X… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-imaging-model-evaluation-benchmark.willie-benchmark
WILLIE Wound Benchmark
Three public wound datasets unified into a single 5-class taxonomy with
fixed splits for classification, segmentation and localization.
The benchmark accompanying WILLIE, published at MLHC 2026.
Developed in the Qian Group, University of Houston.
Models: QianGroup/willie-weights
Code and notebooks: GitHub repository
Paper: MLHC 2026 (link to follow)
What this is
Three public wound datasets — FUSeg, AZH and Medetec — mapped onto one… See the full description on the dataset page: https://huggingface.co/datasets/QianGroup/willie-benchmark.urban-perception-benchmark
Urban Perception Benchmark
Pretty name: Urban Perception Benchmark — Montreal 100Short name: UPB-MTL100License (data): CC BY-NC 4.0 (non-commercial)License (code): MITLanguages: French (source), English (normalized)Modalities: Images + structured annotationsSize: 100 images (50 synthetic, 50 real)Tasks: multi-label and single-choice annotation; evaluation of VLMs on urban perception
This repository hosts the dataset and annotation schema described in the paper:“Do Vision–Language… See the full description on the dataset page: https://huggingface.co/datasets/rsdmu/urban-perception-benchmark.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.doc-openvocab-benchmark
Open-Vocab Document & Figure Classification Benchmark
Given a document or figure image and an arbitrary set of text labels, which one is right? This is a
zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is
scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task
is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label
supervised… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.nra-benchmarks
🧬 NRA Benchmark Datasets
All benchmark datasets for Neural Ready Archive (NRA) — the Rust-native streaming format for ML training.
Train on gigabytes of real data without downloading a single byte. NRA replaces tar.gz and zip for the AI era.
📦 Available Datasets
File
Domain
Source
Files
Size
food-101.nra
🖼️ Vision
ethz/food101
101,000 images
4.7 GB
wikitext.nra
📝 Text
Salesforce/wikitext
23,767 text files
7.6 MB
pokemon.nra
🎨 Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/zevatov/nra-benchmarks.
