datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
geo-benchmarks
im2gps3k and yfcc4k
Two standard image-geolocation test sets, packaged as zips.
File
Contents
im2gps3k.zip
im2gps3k/im2gps3k_places365.csv (columns IMG_ID, AUTHOR, LAT, LON, ...; 2,997 rows) and im2gps3k/images/ (3,000 photos)
yfcc4k.zip
yfcc4k/yfcc4k.csv (columns IMG_ID, OwnerNSID, LAT, LON, ...; 4,536 rows), yfcc4k/yfcc4k.txt (raw YFCC100M metadata) and yfcc4k/images/ (4,536 photos)
Unzip both into one directory to get the layout <root>/{im2gps3k,yfcc4k}/....… See the full description on the dataset page: https://huggingface.co/datasets/kinghorton42/geo-benchmarks.CTTA-AD-Benchmarks
CTTA-AD Benchmarks
Dataset collection for CTTA-AD: Continual Test-Time Adaptation for Unified Few-Shot Visual Anomaly Detection (AAAI 2027 submission).
Datasets
Dataset
Domain
Categories
Train Normal
License
MVTec-AD
Industrial
15
209–391 per category
CC BY-NC-SA 4.0
VisA
Industrial
12
400–905 per category
CC BY-NC-SA 4.0
MVTec-LOCO
Logical
5
varies
CC BY-NC-SA 4.0
BrainMRI
Medical
1
7,500
Research only
LiverCT
Medical
1
1,542
Research only… See the full description on the dataset page: https://huggingface.co/datasets/Hammadhaideerr/CTTA-AD-Benchmarks.nra-benchmarks
🧬 NRA Benchmark Datasets
All benchmark datasets for Neural Ready Archive (NRA) — the Rust-native streaming format for ML training.
Train on gigabytes of real data without downloading a single byte. NRA replaces tar.gz and zip for the AI era.
📦 Available Datasets
File
Domain
Source
Files
Size
food-101.nra
🖼️ Vision
ethz/food101
101,000 images
4.7 GB
wikitext.nra
📝 Text
Salesforce/wikitext
23,767 text files
7.6 MB
pokemon.nra
🎨 Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/zevatov/nra-benchmarks.edge-inference-benchmarks
TinyEdge edge-inference benchmarks
Independently measured latency and accuracy for well-known vision models on
real edge devices (phones, tablets — fleet growing), produced by
TinyEdge, a device cloud for edge-AI benchmarking.
Nothing here is taken from papers or spec sheets: every row is a job executed
on the physical device through TinyEdge's production agent, with accuracy
measured on a fixed 500-image stratified sample of
ImageNet-V2 (matched-frequency)
using a standardized… See the full description on the dataset page: https://huggingface.co/datasets/TinyEdge/edge-inference-benchmarks.
