datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
testdataInter-Edit-Test
Inter-Edit-Test
Official test benchmark release for the CVPR 2026 paper:
Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing
This repository hosts the public release of Inter-Edit-Test, a human-annotated benchmark for the Interactive Instruction-based Image Editing (I^3E) task.
Each sample contains:
a source image,
a coarse user-style interaction mask,
a concise editing instruction,
and a ground-truth edited image.
To simplify large-scale distribution on… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Test.egoxtreme-test
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
📖 Dataset Information
EgoXtreme is a novel large-scale dataset designed for robust egocentric 6D object pose estimation under extreme environmental conditions. It was captured at 30 fps using Aria glasses, providing high-resolution 1408 x 1408 raw fisheye RGB images. The dataset features 15 participants performing diverse interactions with 13 different objects… See the full description on the dataset page: https://huggingface.co/datasets/taegyoun88/egoxtreme-test.DressCode-Testwds_wilds-iwildcam_testwds_country211_test
Country-211 (Test set only)
Original paper: Learning Transferable Visual Models From Natural Language Supervision
Homepage: https://github.com/openai/CLIP/blob/main/data/country211.md
Derived from YFCC100M: https://multimediacommons.wordpress.com/yfcc100m-core-dataset/
Bibtex:
@article{DBLP:journals/corr/abs-2103-00020,
author = {Alec Radford and
Jong Wook Kim and
Chris Hallacy and
Aditya Ramesh and
Gabriel Goh and… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_country211_test.wds_vtab-clevr_closest_object_distance_test
CLEVR Closest Object Distance Webdataset (Test set only)
Original paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Homepage: https://cs.stanford.edu/people/jcjohns/clevr/
Bibtex:
@article{DBLP:journals/corr/JohnsonHMFZG16,
author = {Justin Johnson and
Bharath Hariharan and
Laurens van der Maaten and
Li Fei{-}Fei and
C. Lawrence Zitnick and
Ross B. Girshick}… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_vtab-clevr_closest_object_distance_test.wds_fgvc_aircraft_test
FGVC-Aircraft (Test set only)
Original paper: Fine-Grained Visual Classification of Aircraft
Homepage: https://www.robots.ox.ac.uk/~vgg/data/fgvc-aircraft/
Bibtex:
@techreport{maji13fine-grained,
title = {Fine-Grained Visual Classification of Aircraft},
author = {S. Maji and J. Kannala and E. Rahtu
and M. Blaschko and A. Vedaldi},
year = {2013},
archivePrefix = {arXiv},
eprint = {1306.5151},
primaryClass = "cs-cv"… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_fgvc_aircraft_test.wds_wilds-camelyon17_testwds_wilds-fmow_testGarments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed item… See the full description on the dataset page: https://huggingface.co/datasets/ArtmeScienceLab/Garments2Look-Test-Set-Results.Diff-training-testwds_vtab-clevr_count_all_test
CLEVR Count All Webdataset (Test set only)
Original paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Homepage: https://cs.stanford.edu/people/jcjohns/clevr/
Bibtex:
@article{DBLP:journals/corr/JohnsonHMFZG16,
author = {Justin Johnson and
Bharath Hariharan and
Laurens van der Maaten and
Li Fei{-}Fei and
C. Lawrence Zitnick and
Ross B. Girshick},
title =… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_vtab-clevr_count_all_test.DicFace-test_datasetwds_vtab-dsprites_label_y_position_test
dSprites Y Position (Test set only)
Original paper: beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework
Homepage: https://github.com/deepmind/dsprites-dataset
Bibtex:
@misc{dsprites17,
author = {Loic Matthey and Irina Higgins and Demis Hassabis and Alexander Lerchner},
title = {dSprites: Disentanglement testing Sprites dataset},
howpublished= {https://github.com/deepmind/dsprites-dataset/},
year = "2017",
}
wds_vtab-dsprites_label_x_position_test
dSprites X Position (Test set only)
Original paper: beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework
Homepage: https://github.com/deepmind/dsprites-dataset
Bibtex:
@misc{dsprites17,
author = {Loic Matthey and Irina Higgins and Demis Hassabis and Alexander Lerchner},
title = {dSprites: Disentanglement testing Sprites dataset},
howpublished= {https://github.com/deepmind/dsprites-dataset/},
year = "2017",
}
test-datawds_vtab-dsprites_label_orientation_test
dSprites Orientation (Test set only)
Original paper: beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework
Homepage: https://github.com/deepmind/dsprites-dataset
Bibtex:
@misc{dsprites17,
author = {Loic Matthey and Irina Higgins and Demis Hassabis and Alexander Lerchner},
title = {dSprites: Disentanglement testing Sprites dataset},
howpublished= {https://github.com/deepmind/dsprites-dataset/},
year = "2017",
}
testingtest_fake_imagewds_vtab-dmlab_test
DMLab Frames (Test set only)
Original paper: The Visual Task Adaptation Benchmark
Homepage: https://github.com/google-research/task_adaptation
Bibtex:
@article{zhai2019visual,
title={The Visual Task Adaptation Benchmark},
author={Xiaohua Zhai and Joan Puigcerver and Alexander Kolesnikov and
Pierre Ruyssen and Carlos Riquelme and Mario Lucic and
Josip Djolonga and Andre Susano Pinto and Maxim Neumann and
Alexey Dosovitskiy… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_vtab-dmlab_test.Test_dataset_repo202502_style_disentangle_test_11202502_style_disentangle_test_36test-setGarments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/baajarmah/Garments2Look-Test-Set-Results.Robin-test-dataThis dataset contains the first 40k prompts from LAION/CC/SBU BLIP-Caption Concept-balanced 558K which we use for rapid testing of the Robin model setup on new compute.
This is based on the data used in LLaVA: https://github.com/haotian-liu/LLaVA/blob/main/docs/Data.md
This does about 150 iterations with a batch size of 256 to check checkpointing and final model save.
Objects365_testtest-images202502_style_disentangle_test_20
