datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.cifar100
Dataset Card for CIFAR-100
Dataset Summary
The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).
Supported Tasks and Leaderboards
image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar100.CircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.CiQi-VQA
CiQi-Agent
Github | Model | Dataset | Paper
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Accepted to ECCV 2026
🎯 Overview
CiQi-Agent has been accepted to ECCV 2026.
We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.cityscapesversion https://git-lfs.github.com/spec/v1
oid sha256:4bcf87ecfbbb8e07a01b21415a970c8b53a5283bf6872b657040d3f45c9241f7
size 31
GeoBenchMeta
GeoBench: A Benchmark for Geometric Image Editing
This repository contains the GeoBench benchmark dataset, introduced in the paper Training-Free Diffusion for Geometric Image Editing.
Project Page & Code: https://github.com/CIawevy/FreeFine
GeoBench is designed to evaluate the capability of diffusion models in geometric image editing tasks. It supports various scenarios including object repositioning, reorientation, reshaping, fine-grained partial editing, structure completion… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/GeoBenchMeta.TextPecker-1.5M
TextPecker-1.5M: A Dataset for Training and evaluating TextPecker
This repository contains the TextPecker-1.5M dataset, a new benchmark proposed in the paper "TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering".
Code and Project Page
The official implementation and project details for the TextPecker and TextPecker-1.5M dataset can be found on the GitHub repository:
https://github.com/CIawevy/TextPecker
Sample Usage
You… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/TextPecker-1.5M.cifar100IPL-Cityscapes-Illuminants
IPL-CityscapesIlluminants-dataset
Illuminant modified dataset version of the famous autonomous driving semantic segmentation Cityscapes dataset.
Dataset generation
For each image, we generate a flat (constant) light spectrum and compute the pixel reflectances that obtain the RGM pixel values. Once we have the pixel reflectances, we generate different light spectrums of different dominant wavelengths (colors) and saturations and compute the new modified images. We apply… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/IPL-Cityscapes-Illuminants.cifar-10-cciaa-annual-reports
CIAA Annual Reports — Nepali transcripts, ruled tables and chart data
Machine-readable transcripts of the annual reports of Nepal's Commission for the
Investigation of Abuse of Authority (अख्तियार दुरुपयोग अनुसन्धान आयोग, CIAA) —
all 35 it has published to date. The 1st to 35th reports, fiscal years
BS 2047/48 – 2081/82 (AD 1990–2025).
The CIAA publishes these as PDFs whose text layer is, for several years, legacy
pre-Unicode Devanagari that ordinary extractors turn into… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/ciaa-annual-reports.IPL-Cityscapes-LuminanceContrasts
IPL-CityscapesLuminanceContrasts-dataset
Controled luminance and contrasts modified dataset version of the famous autonomous driving semantic segmentation Cityscapes dataset.
Dataset generation
For each original image, we convert it to ATD color space. Once in this space, we compute its mean luminance, achromatic contrast and chromatic contast. We modify each of its characteristics in turns from 0.5 to 1.5 of its original value. Then we return the image to RGB space. We… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/IPL-Cityscapes-LuminanceContrasts.cif-dataset
Cracks in the Foundation
A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories:
Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling
Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one.
Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples.
Splits
Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.cifar10cinic10
Dataset Card for CINIC-10
CINIC-10 has a total of 270,000 images equally split amongst three subsets: train, validate, and test. This means that CINIC-10 has 4.5 times as many samples than CIFAR-10.
Dataset Details
In each subset (90,000 images), there are ten classes (identical to CIFAR-10 classes). There are 9000 images per class per subset. Using the suggested data split (an equal three-way split), CINIC-10 has 1.8 times as many training samples as in CIFAR-10.… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/cinic10.civitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/civitai-top-nsfw-images-with-metadata.cityscapesM-BEIR-CIRR-ImagesNova
NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI
An open-world generalization benchmark under clinical distribution shift
Dataset on 🤗 Hugging FaceFor academic, non-commercial use only
🔖 Citation
If you find this dataset useful in your work, please consider citing it:
@article{bercea2025nova,
title={NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI},
author={Bercea, Cosmin I. and Li, Jun and… See the full description on the dataset page: https://huggingface.co/datasets/c-i-ber/Nova.MMSearch-Plus
MMSearch-Plus✨: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
Official repository for the paper "MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents".
🌟 For more details, please refer to the project page with examples: https://mmsearch-plus.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface Dataset] [🏆 Leaderboard]
💥 News
[2025.09.26] 🔥 We update the arXiv paperand release all MMSearch-Plus data samples in… See the full description on the dataset page: https://huggingface.co/datasets/Cie1/MMSearch-Plus.cibench
cibench — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror (config cibench_gen). Source open-compass/CIBench (GitHub); data zip in open-compass/opencompass releases — license Apache-2.0, unchanged; all rights remain with the original authors.
cifar100_2_to_100_constant_size_dataset
Dataset Card for "cifar100_2_to_100_constant_size_dataset"
More Information needed
cinematic-stillscifar100-enrichedThe CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).cityscape-adverse
Cityscape‑Adverse
A benchmark for evaluating semantic segmentation robustness under realistic adverse conditions.
Overview
Cityscape‑Adverse extends the original Cityscapes dataset by applying eight realistic environmental modifications—rainy, foggy, spring, autumn, snowy, sunny, night, and dawn—using diffusion‑based image editing. All transformations preserve the original 2048×1024 semantic labels, enabling direct evaluation of model robustness in… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cityscape-adverse.floorplans-cityscapes
Dataset Summary
This is a curated collection of floorplan images sourced from across the internet. It is intended for research in architectural AI, layout generation, and urban scene understanding.
Data format: Image files with associated integer labels.
Sources: Publicly available images from various web sources (This dataset is one unified collections).
Purpose: Educational and research use.
Dataset Structure
The dataset follows the standard Hugging Face Image… See the full description on the dataset page: https://huggingface.co/datasets/wheres-my-python/floorplans-cityscapes.mmmu-ci-50TTA-Cityscapes-Cwds_vtab-cifar10rl-game-traces-civilization-6
文明6
This public dataset repository contains gameplay trace data uploaded from F:\文明6.
Contents
Files: 413
Total local size: 363.61 GB
Generated: 2026-06-09T09:17:06+00:00
File Types
.jsonl: 123
.json: 106
.png: 90
.parquet: 32
.txt: 31
.mkv: 30
.xlsx: 1
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage, audio, and asset… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-civilization-6.
