datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MARS-Hyperspectral-EnMAP-PRISMA-v2025
MARS-Hyperspectral dataset (EnMAP and PRISMA) - version v2025
Updated version of the dataset. More information will be added soon.
MARS-Hyperspectral
MARS-Hyperspectral dataset
This repository contains the data for the work publicly presented at:
[ArXiv preprint] Růžička, Mateo-García, Irakulis-Loitxate et al., Operational machine learning for remote spectroscopic detection of CH4 point sources, arXiv preprint arXiv:2511.07719 (2025).
[Poster] Růžička, Mateo-García, Irakulis-Loitxate et al., Machine Learning Models for Multi-sensor Detection of Methane Leaks in Hyperspectral Data, June. 23-27, 2025, ESA Living Planet… See the full description on the dataset page: https://huggingface.co/datasets/UNEP-IMEO/MARS-Hyperspectral.Hypersim-Processedhyperspectral-orchard
Living Optics Orchard Dataset
Overview
This dataset contains 435 images of captured in one of the UK's largest orchards, using the Living Optics Camera.
The data consists of RGB images, sparse spectral samples and instance segmentation masks.
The dataset is derived from 44 unique raw files corresponding to 435 frames.
Therefore, multiple frames could originate from the same raw file.
This structure emphasized the need for a split strategy that avoided data leakage.
To… See the full description on the dataset page: https://huggingface.co/datasets/LivingOptics/hyperspectral-orchard.hypersim-exampleshypersim_relative_depth
Hypersim Relative Depth
This repository contains a repackaged and preprocessed version of the
Hypersim Dataset, prepared for relative
depth training with the Marigold V2 codebase.
The repository contains RGB/depth pairs and split metadata organized into
downloadable tar archives. The preprocessing and archive layout are intended
for use with the Marigold V2 data-loading pipeline.
Source dataset
The data originates from:
Dataset: Hypersim
Original repository:… See the full description on the dataset page: https://huggingface.co/datasets/obukhovai/hypersim_relative_depth.hypersimMSI_productsHyperKvasir
Dataset Card for HyperKvasir
HyperKvasir is the largest publicly available dataset for gastrointestinal (GI) endoscopy, consisting of over 110,000 images and 374 videos, including labeled, unlabeled, and segmented samples. It supports tasks such as classification, segmentation, object detection, and anomaly detection in GI disease diagnostics.
Dataset Details
Dataset Description
HyperKvasir is a comprehensive multi-class dataset collected from… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet-HOST/HyperKvasir.OS-Atlas_ScreenSpotPACE_productshyperspectral-fruit
Living Optics Hyperspectral Fruit Dataset
Overview
This dataset contains 100 images of various fruits and vegetables captured under controlled lighting, with the Living Optics Camera.
The data consists of RGB images, sparse spectral samples and instance segmentation masks.
From the 100 images, we extract >430,000 spectral samples, of which >85,000 belong to one of the 19 classes in the dataset. The rest of the spectra can be used for negative sampling when training… See the full description on the dataset page: https://huggingface.co/datasets/LivingOptics/hyperspectral-fruit.layout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation.
Project page: https://orangesodahub.github.io/SceneCraft
Code: https://github.com/OrangeSodahub/SceneCraft
EMIT_productsamazon-berkeley-objects
Amazon Berkeley Objects
This is a Hugging Face metadata mirror of the Amazon Berkeley Objects dataset
for reproducible research and HyperView demos. The original dataset is provided
by Amazon.com and UC Berkeley.
This mirror stores metadata tables and official S3 asset URLs. It does not
duplicate catalog images, turntable images, or 3D models as binary files.
Load
from datasets import load_dataset
listings = load_dataset("hyper3labs/amazon-berkeley-objects"… See the full description on the dataset page: https://huggingface.co/datasets/hyper3labs/amazon-berkeley-objects.hypernet-image-to-3d-data
Hypernet Image-to-3D Dataset
Companion dataset to the hypernet and image-to-3d-deepsdf repositories.
Contents
Path
Size
Description
manifest.json
<1 MB
obj_idx ↔ uid ↔ LVIS category mapping for 1000 shapes
views.json
<1 KB
64 Fibonacci-sphere camera poses
watertight/
~21 GB
985 watertight .obj meshes (filename = uid hash)
sdf_samples/
~3 GB
976 .npz files, each with 200K (point, sdf) pairs
multiview/
~1.3 GB
976 .tar files; each tarball contains 64 PNG… See the full description on the dataset page: https://huggingface.co/datasets/bobthebuilderinternational/hypernet-image-to-3d-data.hyper-scenery-sketches-dataset
Shiro's Hyper-Scenery Dataset
This dataset contains 1000 synthetic image-text pairs used to train the Hyper-Brain.
Images: High-detail sceneries and sketches.
Prompts: Detailed 8K descriptive text for each image.
hypernet-checkpoints
Hypernetwork → Shape pipeline checkpoints
Trained models and processed data for image-to-3D experiments documented at
BOB-THE-BUILDER-in/Hypernetwork.
Contents
tier_essential.tar.gz (1.98 GB) — anchors, autoencoder, mappers, results
tier_data.tar.gz (2.67 GB) — watertight meshes, SDF samples, image-SIRENs, shape-SIRENs
tier_hypernets.tar.gz (6.69 GB) — 100 trained hypernets (one per training shape)
Source
Mesh data derived from Objaverse 1.0.
Individual… See the full description on the dataset page: https://huggingface.co/datasets/bobthebuilderinternational/hypernet-checkpoints.Hypersim_600_resize_float32hyper-kvasir-labeled-imagesHyperKvasir
Labeled images In total, the dataset contains 10,662 labeled images stored using the JPEG format. The images can be found in the images folder. The classes, which each of the images belongto, correspond to the folder they are stored in (e.g., the ’polyp’ folder contains all polyp images, the ’barretts’ folder contains all images of Barrett’s esophagus, etc.). The number of images per class are not balanced, which is a general challenge in the medical field due to the fact that some… See the full description on the dataset page: https://huggingface.co/datasets/sahilur/hyper-kvasir-labeled-images.hyper_drive
Towards automated analysis of large environments, hyperspectral sensors must be adapted into a format where they can be operated from mobile robots. In this dataset, we highlight hyperspectral datacubes collected from the Hyper-Drive imaging system. Our system collects and registers datacubes spanning the visible to shortwave infrared (660-1700 nm) in 33 wavelength channels. The system also simultaneously captures the ambient solar spectrum reflected off a white reference tile. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nhanson2/hyper_drive.HyperNeRF-Annotation
This is the language query annotations for the HyperNeRF dataset, which are used in 4DLangSplat For original dataset, please visit https://github.com/google/hypernerf. For the usage of annotations, please visit https://github.com/zrporz/4DLangSplat
license: cc-by-nc-4.0
Hyperphantasia
A Benchmark for Evaluating the
Mental Visualization Capabilities of Multimodal LLMs
Mohammad Shahab Sepehri
Berk Tinaz
Zalan Fabian
Mahdi Soltanolkotabi
Github Repository
Hyperphantasia is a synthetic Visual Question Answering (VQA) benchmark dataset that probes the mental visualization capabilities of Multimodal Large Language Models (MLLMs) from a vision perspective. We reveal that state-of-the-art models struggle with simple tasks that require visual… See the full description on the dataset page: https://huggingface.co/datasets/shahab7899/Hyperphantasia.hyperlog-ig-assetshypershadow
HyperShadow
Paper: arXiv:2607.14419
A benchmark for one question: given a 3D point cloud, can you tell whether
it is an ordinary 3D object or the 3D projection (the "shadow") of an
object from a higher spatial dimension?
Most datasets that say "4D" mean 3D plus time. Here the extra dimensions
are spatial: label 1 clouds are projections of objects living in R^4, R^5
or R^6 (hyperspheres, tesseracts, Clifford tori, duocylinders, hypertori,
random smooth manifolds), rotated… See the full description on the dataset page: https://huggingface.co/datasets/AkshaySasi/hypershadow.salmonella-serovar-hyperspectral
Salmonella Serovar Hyperspectral Microscopy (Foods 2025)
Salmonella Serovar Hyperspectral Microscopy is an image dataset for foodborne bacterial classification using hyperspectral imaging. It was created to support research in rapid pathogen identification, enabling models to classify Salmonella serovars directly from microscopy images without the need for selective enrichment.
Companion spectral dataset: The single-cell spectral features (tabular, 25,972 rows) extracted from these… See the full description on the dataset page: https://huggingface.co/datasets/food-ai-nexus/salmonella-serovar-hyperspectral.hyperspectral-fruit
Living Optics Hyperspectral Fruit Dataset
Overview
This dataset contains 100 images of various fruits and vegetables captured under controlled lighting, with the Living Optics Camera.
The data consists of RGB images, sparse spectral samples and instance segmentation masks.
From the 100 images, we extract >430,000 spectral samples, of which >85,000 belong to one of the 19 classes in the dataset. The rest of the spectra can be used for negative sampling when… See the full description on the dataset page: https://huggingface.co/datasets/duc08042006/hyperspectral-fruit.synthetic-diabetes-hypertension-NCD-screening-WHO-HEARTS
Synthetic Diabetes & Hypertension NCD Screening Dataset (Adults 18-80) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/synthetic-diabetes-hypertension-NCD-screening-WHO-HEARTS.africa-synth-diabetes-ncd-diabetes-hypertension-all
Synthetic Diabetes & Hypertension NCD Screening Dataset (Adults 18-80) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-diabetes-ncd-diabetes-hypertension-all.hyperscale-datacenter-segmentation-naip
Hyperscale Data Center Segmentation (NAIP)
Hand-digitized training data for detecting hyperscale data center
footprints in aerial imagery, with a trained baseline model.
Contents
path
what it is
datacenters.geojson
190 hand-digitized data center footprint polygons (QGIS; named facilities, e.g. vantage_0)
chips/
757 NAIP aerial chips, 256×256 px, 4-band RGBN, 0.6 m resolution (GeoTIFF, georeferenced) — labeled facilities plus surrounding negatives… See the full description on the dataset page: https://huggingface.co/datasets/rbhughes/hyperscale-datacenter-segmentation-naip.
