datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RobotDesign1M
RobotDesign1M: A Large-scale Dataset for Robot Design Understanding
RobotDesign1M is a large-scale, multimodal dataset for robot design understanding, built from image–text data curated from scientific literature across a wide range of robotics domains. It is designed to support research on design-aware foundation models, including design image generation, visual question answering about designs, and design image retrieval.
📄 Paper: RobotDesign1M: A Large-scale Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/RobotDesign1M.cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.GenEvolve-Data-Bench
GenEvolve Data and Bench
This repository contains the open-source data release for GenEvolve:
Config
Directory
Records
Images
Purpose
sft
GenEvolve-Data-SFT/
9,000 trajectories
50,291 reference images
supervised cold-start trajectories
rl
GenEvolve-Data-RL/
3,175 prompts
3,175 GT images
self-evolution / RL training prompts
bench
GenEvolve-Bench/
594 prompts
594 GT images
held-out evaluation benchmarkAll metadata is provided in both JSONL and Parquet. The Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench.ai4g-flood-dataset
Flood Detection Dataset
Introduction
This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide.
Key features:
Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.MathVista
Dataset Card for MathVista
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Data Visualization
Data Source
Automatic Evaluation
License
Citation
Dataset Description
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts. It consists of three newly created datasets, IQTest, FunctionQA, and PaperQA, which address the missing visual domains and are tailored to evaluate logical… See the full description on the dataset page: https://huggingface.co/datasets/AI4Math/MathVista.dental-panoramic-xray-yolo
Dental Panoramic X-Ray Detection Dataset (YOLO Format)
Combined dataset for dental pathology detection on panoramic radiographs, in YOLO format. Built for training liodon-ai/dental-panoramic-detector.
Classes
ID
Name
Description
0
caries
Dental caries and deep caries
1
periapical_lesion
Periapical / apical periodontitis
2
impacted_tooth
Impacted and wisdom teeth
Dataset Sources
Source
Images
Boxes
License
DENTEX
724
3… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/dental-panoramic-xray-yolo.ViFailback-Dataset
ViFailback Dataset: Real-World Robotic Manipulation Failure Dataset with Visual Symbol Guidance
A real-world dataset for diagnosing, correcting, and learning from robotic manipulation failures via visual symbols.
ViFailback is a large-scale, real-world robotic manipulation failure dataset introduced in the CVPR 2026 paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols". It introduces visual… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/ViFailback-Dataset.ai2d@misc{kembhavi2016diagram,
title={A Diagram Is Worth A Dozen Images},
author={Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Minjoon Seo and Hannaneh Hajishirzi and Ali Farhadi},
year={2016},
eprint={1603.07396},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
polyu-storyworld-characterscdn
Nunchaku CDN
Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.OVMM_objectsNYC-CDIndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.jev-ai-api-guide-assetsaitw-processed-labeled-full
AiTW Processed Full with App Labels
This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset.
Why This Exists
AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.BMD-45
BMD-45: Bengaluru Mobility Dataset
A large-scale CCTV vehicle detection benchmark for Indian urban traffic
Dataset Summary
BMD-45 is a large-scale, India-specific vehicle detection dataset captured from 3,679 operational CCTV cameras across Bengaluru — one of the world's most traffic-congested megacities.
Statistic
Value
Total images
45,986 (1920×1080 RGB)
Total annotations
≈ 481,947 bounding boxes
Vehicle classes
14 fine-grained categories
Camera… See the full description on the dataset page: https://huggingface.co/datasets/iisc-aim/BMD-45.BlueMO
BlueMO
🚀 BlueMO: A Comprehensive Collection of Challenging Mathematical Olympiad Problems from the Little Blue Book Series
BlueMO is a comprehensive and challenging dataset comprising mathematical olympiad problems paired with detailed solutions, meticulously curated from the esteemed "Little Blue Book" (小蓝书) series (Second Edition)—a vital resource for Chinese students training for national and international olympiad math competitions.Designed to advance and… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/BlueMO.mid-space
MID-Space: Aligning Diverse Communities’ Needs to Inclusive Public Spaces
A new version of the dataset will be released soon, incorporating user identity markers and expanded annotations.
LIVS PAPER
Click below to see more:
Overview
The MID-Space dataset is designed to align AI-generated visualizations of urban public spaces with the preferences of diverse and marginalized communities in Montreal. It includes textual prompts, Stable Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/mila-ai4h/mid-space.AID
Aerial Image Dataset (AID)
Description
The Aerial Image Dataset (AID) is a scene classification dataset consisting of 10,000 RGB images, each with a resolution of 600x600 pixels. These images have been extracted using Google Earth and cover various scenes from regions and countries around the world. AID comprises 30 different scene categories, with several hundred images per class.
The new dataset is made up of the following 30 aerial scene types: airport, bare… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/AID.airbot-fold-cloth-mcapVisionArena-Chat
VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes
200k single and multi-turn chats between users and VLM's collected on Chatbot Arena.
WARNING: Images may contain inappropriate content.
Dataset Details
200K conversations
45 VLM's
138 languages
~43k unique images
Question Category Tags (Captioning, OCR, Entity Recognition, Coding, Homework, Diagram, Humor, Creative Writing, Refusal)
Dataset Description
200,000… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/VisionArena-Chat.RAIDThis dataset is for testing the adversarial robustness of AI-Generated Image Detectors, as described in the paper RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors.
olympiadbenchAIGIBench
Is Artificial Intelligence Generated Image Detection a Solved Problem?
Ziqiang Li1, Jiazhen Yan1, Ziwen He1, Kai Zeng2, Weiwei Jiang1, Lizhi Xiong1, Zhangjie Fu1‡
‡Corresponding author
1Nanjing University of Information Science and Technology 2University of Siena
Paper | GitHub Repository
This repository is the official dataset of the AIGIBench.
AIGIBench dataset contains two types of training and 25 test subsets. This dataset has the following advantages:
Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/HorizonTEL/AIGIBench.aitod-v2
AI-TOD-v2
AI-TOD-v2, the tiny-object detection benchmark in aerial images, packed once with the official v2 annotations kept whole, so it loads in one line and no data path has to be configured:
from datasets import load_dataset
ds = load_dataset("shijli/aitod-v2") # 11214 train / 2804 validation / 14018 test
AI-TOD cuts 28036 images of 800 x 800 pixels from xView, DOTA-v1.5, VisDrone2018-Det, Airbus Ship Detection and DIOR, and annotates eight classes whose mean object size… See the full description on the dataset page: https://huggingface.co/datasets/shijli/aitod-v2.anything-v3.0-glazed
Dataset Card for Anything v3.0 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by Linaqruf/anything-v3.0
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/anything-v3.0-glazed.AirGoal-10k
AirGoal-10k
AirGoal-10k is an aerial image-goal navigation dataset released with
UA-NWM: Uncertainty-Aware World Model for Aerial Image-Goal Navigation.
Project page: https://duryi.github.io/UA-NWM-Project-Page/Code: https://github.com/DurYi/UA-NWMPaper: https://arxiv.org/abs/2608.05597
Dataset Summary
AirGoal-10k contains 11,000 aerial navigation trajectories for image-goal navigation. Each trajectory
contains 12 RGB observations and trajectory metadata. The test… See the full description on the dataset page: https://huggingface.co/datasets/DurYi/AirGoal-10k.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
The dataset was initially developed exclusively for military aircraft detection, but was later expanded to include commercial airliners for a broader and more challenging detection task.
The dataset contains 103 military aircraft types and 11 commercial airliner types.
Military aircraft: A10, A400M, AG600, AH64, AKINCI, AV8B, An124, An22, An225, An72, B1, B2, B21, B52, Be200, C1… See the full description on the dataset page: https://huggingface.co/datasets/a2015003713/military-aircraft-detection-dataset.
