timm/imet-open-access
iMet Open Access A multi-label image classification dataset of 259,559 artworks and objects from The Metropolitan Museum of Art, built from the Met's official CC0 Open Access release (metmuseum/openaccess). It follows the spirit of the Kaggle iMet Collection challenges (FGVC6 2019 / FGVC7 2020): long-tail, fine-grained attribute recognition from museum photography, with subject tags, culture, medium (material / technique) and country labels. This is not the Kaggle iMet data.… See the full description on the dataset page: https://huggingface.co/datasets/timm/imet-open-access.
iMet Open Access
A multi-label image classification dataset of 259,559 artworks and objects from The Metropolitan Museum of Art, built from the Met's official CC0 Open Access release (metmuseum/openaccess). It follows the spirit of the Kaggle iMet Collection challenges (FGVC6 2019 / FGVC7 2020): long-tail, fine-grained attribute recognition from museum photography, with subject tags, culture, medium (material / technique) and country labels.
This is not the Kaggle iMet data. Those releases can't be redistributed and their test labels were never published. This dataset is an independent, openly licensed rebuild from the same source (The Met's Open Access catalog and images). It uses its own label normalization, vocabulary and splits.
Splits
- 10 folds, assigned by multi-label iterative stratification (Sechidis et al., 2011; seed 42) over groups on the
labelsvocabulary. train = folds 0–7, validation = fold 8, test = fold 9. Thefoldcolumn allows k-fold use. - Near-duplicate grouping: objects sharing the same photograph (identical perceptual hashes), and DINOv2 mutual nearest neighbours with cosine similarity ≥ 0.95 (the same design or series, e.g. several impressions of one print, pairs from a furniture set), are kept in the same fold.
group_id/group_sizerecord this. - Every
labelsclass is present in validation and test. - Excluded: 312 source rows without an image and 3 blank images.
Label sets
All label columns are lists of ClassLabel indices, so they can be used directly as multi-label targets.
The combined sets use type::label names (tag::Horses, culture::japanese, medium::porcelain, country::egypt), like iMet's labels.csv. A missing label does not mean a negative: tags in particular are incomplete (see Limitations).
How the labels were made
All labels come from the Met's own catalog fields. Normalization was LLM-assisted and every decision was reviewed; the maps and decision tables are available on request.
- Tags: the Met's curated subject keyword tags (Getty AAT / Wikidata linked). Equivalent names are merged into one label, e.g. Greek/Roman gods (
Aphrodite / Venus,Eros / Cupid) and synonyms (Leopards / Panthers). A few exhibition labels are dropped. - Culture: the free-text
culturefield is normalized to consistent demonyms (Japan/Japanese →japanese), withparent/sublabels for regional schools (greek+greek/attic). - Historical cultures stay distinct:
sasanian,coptic,etruscan, … - Pre-Columbian objects get culture-area labels (
mesoamerican,andean). - Iran is split into
iranian(pre-Islamic) andpersian(Islamic era). - Uncertain attributions ("American or European") become a compound label (
american or european), plus the shared parent when there is one. The individual alternatives are never assigned. - When
cultureis empty, it is recovered from other catalog fields.culture_sourcerecords which:
- Medium: the
mediumtext is split into material / technique terms and normalized: spelling and synonyms merged, print states, formats and tools dropped. Terms map to a ≤ 2-level hierarchy (hard-paste porcelain→ alsoporcelain). - Country: the Met
countryvalue, used only when itsgeographyTypestates origin, production or find-spot (not "Published in", "Depicted", etc.).
Images
Primary image of each object, re-encoded from the full-resolution Open Access JPEGs:
- Size: shorter side 768 px, longer side capped at 1920 px, aspect ratio preserved. Images are never upscaled.
- Resampling: full decode, then Pillow LANCZOS.
- Colour: embedded ICC profiles converted to sRGB; grayscale, CMYK and palette images converted to RGB.
- Orientation and frames: EXIF orientation applied; first frame of multi-frame files.
- Format: JPEG quality 95, 4:2:0.
width/heightgive the encoded size andorig_width/orig_heightthe source size.
Other columns
object_id, accession_number, object_url (Met collection page), title, object_name, classification_raw, is_highlight, culture_raw, culture_source, culture_qualifier, medium_raw, dimensions_raw, height_cm / width_cm / depth_cm (parsed from dimensions_raw; null when not parseable), period, dynasty, object_date, object_begin_date / object_end_date, country_raw, geography_type, artist_display_name, artist_nationality, tags_raw / tag_aat_urls / tag_wikidata_urls, group_id, group_size.
Usage
python train.py --dataset hfds/timm/imet-open-access --train-split train --val-split validation \
--task multilabel --target-key labels --num-classes 1302 \
--model convnext_small --pretrained --img-size 384 --batch-size 64 --epochs 30 --ampAny other label column works as --target-key (e.g. --target-key medium_labels --num-classes 387).
from datasets import load_dataset
ds = load_dataset('timm/imet-open-access', split='validation')
names = ds.features['labels'].feature.names
print([names[i] for i in ds[0]['labels']])Limitations
- Tags are incomplete: about half the objects have any tag, and those that do are not exhaustively tagged. A missing tag isn't a negative, so losses robust to missing labels and per-type evaluation are recommended.
- Culture is a curatorial attribution, not ground truth. The
artistsource uses artist nationality as a proxy for the object's culture. Compound "X or Y" labels mean uncertain attribution, so an evaluation can give partial credit for predicting either component. - Coverage is uneven: most European decorative arts (laces, textiles, silver) have no culture label, and country is only recorded for about 21% of objects. Department mix and labels are correlated (e.g. prints dominate tags such as
PortraitsorActresses). - Many classes have only a few validation / test examples. Report macro and micro mAP, and consider per-type metrics.
License and attribution
The Met's Open Access images and data are released under CC0 1.0 (public domain dedication), and this derived dataset (re-encoded images, normalized labels, splits) is released under CC0 1.0 as well. As The Met requests, please credit The Metropolitan Museum of Art, Open Access and link to the object pages (object_url) where practical.
Background: Zhang et al., The iMet Collection 2019 Challenge Dataset, arXiv:1906.00901.
