datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TerraMesh
TerraMesh
A planetary‑scale, multimodal analysis‑ready dataset for Earth‑Observation foundation models: TerraMesh merges data from Sentinel‑1 SAR, Sentinel‑2 optical, Copernicus DEM, NDVI, and land‑cover sources into more than 9 million co‑registered patches ready for large‑scale representation learning.
You find more information about the data sampling and preprocessing in our paper: TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data.
Samples from the TerraMesh… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh.Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.CoordBench
CoordBench
A unified benchmark suite for evaluating location encoders such as SatCLIP, GeoCLIP, Climplicit, and MIND.
The dataset contains 40 normalized source tables from 13 source families. The paper's evaluation suite
uses 52 datasets and 78 prediction targets drawn from this mirror. The source files previously lived across GitHub,
figshare, GCS, Socrata, Zenodo, and Google Drive.
Intended use
Use the normalized tables to compare coordinate-to-embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/CoordBench.MINDSET
MINDSET
MINDSET is the pretraining dataset for MIND, a coordinate-only location encoder distilled from static location encoder teachers and annual AlphaEarth Foundations (AEF) embeddings.
We release the embeddings at the 12.1M training coordinates. The dataset contains 12,099,072 land coordinates in WGS84. Coordinates are dense around cities and not uniformly sampled over land.
The files are in GeoParquet format and can be joined on point_id:
file
grain
rows
columns… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/MINDSET.TerraMesh-Masks
TerraMesh-Masks
TerraMesh-Masks is a dataset for open-vocabulary segmentation of satellite imagery. This dataset provides binary segmentation masks with captions that extend the samples from TerraMesh.
We also provide an human-verfied evaluation benchmark, called TerraMesh-Masks-Eval.
Examples from the training subset:
Usage
Download the data loading code from GitHub and install requirements with pip install -r requirements.txt. For development, you can… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks.pad-pid-geospatial
Geospatial data use in World Bank PADs and PIDs
The data behind How often do World Bank projects use geospatial data?.
Snapshot 938866a81eab, annotations to 2026-09-25 21:57 UTC. The page shows the same snapshot id.
Headline
Out of: those that use data
Out of: all
Projects
1,165 of 1,932 (60.3%)
1,165 of 2,261 (51.5%)
Documents
1,637 of 3,544 (46.2%)
1,637 of 4,494 (36.4%)
A project uses geospatial data when at least one data mention in the… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/pad-pid-geospatial.TerraMesh-Masks-Eval
TerraMesh-Masks-Eval
TerraMesh-Masks-Eval is a human-verified benchmark dataset to evaluate open-vocabulary segmentation models on satellite imagery. This dataset provides binary segmentation masks with captions togehther with input samples from TerraMesh.
We also provide a training dataset, called TerraMesh-Masks.
Examples from the evaluation subset:
Usage
Download the data loading code from GitHub and install requirements with pip install -r requirements.txt.… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks-Eval.brazil-wildfire-geospatial-dataset
Banco Histórico de Incêndios no Brasil — 2018–2025
Banco de dados com 1,43 milhão de focos de calor registrados no Brasil entre 2018 e 2025, enriquecidos com dados meteorológicos (ERA5), cobertura do solo (MapBiomas) e altitude (SRTM).
Construído a partir de fontes públicas oficiais para suporte a pesquisas científicas sobre incêndios florestais.
Tabelas disponíveis
Tabela
Arquivos
Linhas
Descrição
focos_analise
focos_analise/*.parquet
1.430.756
Tabela… See the full description on the dataset page: https://huggingface.co/datasets/mateus-pcosta/brazil-wildfire-geospatial-dataset.florida_geospatialabdullahkhan70_global-street-food-3-continent-geospatial-index
Global Street Food: 3 Continent Geospatial Index
3 Countries street food index: GPS, local prices, and hygiene ratings
Dataset Info
Source: Kaggle
Original Size: 0.11 MB
Kaggle Downloads: 26
Files: 3
Files
mexico_street_food_vendor.csv
pakistan_street_food_vendor.csv
thailand_street_food_vendor.csv
Mirrored from Kaggle
US_GeoSpatial_Dataset_by_HNM
US_GeoSpatial_dataset_by_HNM
Dataset Description
This dataset contains 56 records with 16 features.
Dataset Summary
Metric
Value
Total Rows
56
Total Columns
16
Numeric Columns
10
Categorical Columns
6
Missing Values
0 (0.00%)
Duplicate Rows
0
Memory Usage
19.51 MB
Dataset Structure
Data Fields
Column
Type
Sample/Range
Unique Values
Missing %
geo_id
int64
Range: [1.00, 78.00]
56
0.0%… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/US_GeoSpatial_Dataset_by_HNM.
