NoeFlandre/osm-polygon-selection
osm-polygon-selection dataset A curated set of OpenStreetMap polygons from 310 geographic units — sovereign countries plus sub-country regions like Brazilian states, Chinese provinces, Indian zones, US states, Canadian provinces, Japanese regions, and Indonesian islands — classified by size bin (small / medium / large, area in [0.1, 100] km²) and tagged by continent (Natural Earth admin0 lookup). Size bins: small — area in [0.1, 1) km² (10,000 m² to 1 km², roughly 100 m × 100 m… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-selection.
osm-polygon-selection dataset
A curated set of OpenStreetMap polygons from 310 geographic units — sovereign countries plus sub-country regions like Brazilian states, Chinese provinces, Indian zones, US states, Canadian provinces, Japanese regions, and Indonesian islands — classified by size bin (small / medium / large, area in [0.1, 100] km²) and tagged by continent (Natural Earth admin0 lookup).
Size bins:
- `small` — area in [0.1, 1) km² (10,000 m² to 1 km², roughly 100 m × 100 m to ~1 km × 1 km). Examples: a city block, a small park, a single farm field, a small wood lot, a residential courtyard, a parking lot, an industrial yard.
- `medium` — area in [1, 10) km² (1 km² to 10 km², roughly 1 km × 1 km to 3 km × 3 km). Examples: a large park, a small village/town footprint, a reservoir, a forest patch, an industrial zone, a golf course, a cemetery, a nature reserve.
- `large` — area in [10, 100] km² (10 km² to 100 km², roughly 3 km × 3 km to 10 km × 10 km). Examples: a large forest, a big lake, an entire town or small city, a large military training area, a national park section, a sizable agricultural region.
Polygons smaller than 0.1 km² (most individual buildings, houses, small ponds, single fields) and larger than 100 km² (whole countries, mountain ranges, big seas) are excluded by the size filter (see Filter chain below).
Status: All 310 geographic units are extracted end-to-end.
Total polygons: 16,297,690 (combined parquet: combined/all_world.parquet).
Coverage
This dataset processes one parquet per Geofabrik PBF region. Each region is bucketed in per_country/ as either a sovereign country or a sub-country region (state, province, federal district, island group, zone):
Total: 310 geographic units across all 6 inhabited continents plus Oceania. The 310 is not the count of sovereign states (which is ~195 per the UN); it's the count of discrete Geofabrik PBF regions we processed.
Layout
This dataset is split across five subfolders so you can pull only what you need:
Start with sample/ or preview/ for a quick look. Pull per_country/<country>/<country>.parquet for a single-country study. Use combined/all_world.parquet for cross-country work or splits/<train/val/test>.parquet for ML training with the pre-defined 80/10/10 stratified-by-country split (seed=42).
What's in this dataset
Each row is one OSM polygon (closed way or multipolygon relation) that passed our filter chain (see below). The polygon geometry itself is included in the row as WKT (or WKB if OSM_POLYGON_GEOMETRY=wkb is set when the dataset is built) so you can render, query, or reproject it directly without re-deriving from centroid+area.
Provenance
- Pipeline version: v0.1.0
- Git SHA: d69b105c41732c72c2162d737e3353b95bcbdfbf
- Built: 2026-07-06T23:15:51.406227
- Source: Geofabrik regional extracts (
https://download.geofabrik.de/) - Whitelist: 22,075 OSM
key=valuetags from osm-stats (seedocs/whitelist_decisions.mdin the project repo, or read the full rationale in the blog post). The whitelist is designed to filter polygons by landuse-style tags (natural,landuse,leisure,amenity, etc.) so the dataset focuses on physical land-cover / land-use features rather than buildings, addresses, or points of interest.
Geographic distribution
(Each circle is one polygon from the sample/ folder, color-coded by country. Circle size is proportional to sqrt(area_km2).)
Size-bin distribution (full dataset)
Counts every polygon in the 16,297,690-polygon dataset by size_bin, computed directly from combined/all_world.parquet via pyarrow.compute.value_counts. Percentages are exact ratios over the entire dataset, not a sample.
Example row
Here is one concrete row from the Liechtenstein parquet file (a natural=* polygon, fully filled-in with all 13 columns):
This row is representative: the full-dataset distribution above shows ~80% small, ~18% medium, ~2% large, and the dominant whitelist tag families (natural=*, landuse=*, leisure=*) account for the majority of matched_tag values.
Filter chain
Each polygon in this dataset has passed three filters:
- Size filter (Stage 0): area in [0.1, 100] km². Polygons smaller than 0.1 km² or larger than 100 km² are dropped.
- Whitelist filter (Stage 2): at least one OSM tag in the 22,075-tag whitelist. The whitelist is derived from a clustering of OSM tags across both
tfidfandembeddingsanalyses. - Classify (Stage 3): continent assigned via Natural Earth admin0 shapefile, size_bin assigned by area.
Train / val / test split
Every row in every parquet (per_country/<country>/<country>.parquet and combined/all_world.parquet) carries a `split` column with one of three values: train, val, or test.
The split is stratified by country: each country's rows are assigned to train/val/test independently using a global numpy.random.default_rng seeded once with 42 and offset per country. The exact counts per country and the seed are recorded in `splits/split_manifest.json`.
To load only one split (e.g. for training), filter in pyarrow:
import pyarrow.compute as pc
import pyarrow.parquet as pq
table = pq.read_table("combined/all_world.parquet")
train = table.filter(pc.equal(table["split"], "train"))The split is deterministic and re-runnable:
uv run python scripts/make_split.py # default seed=42, 80/10/10
uv run python scripts/make_split.py --seed 7 # different reproducible splitPer-country summary
Back to the dataset root
