buildchange/segmentation_training_data
Build Change augmented segmentation dataset Paired images and segmentation masks, including augmented (rotated, flipped, transformed) copies of each original photo, paired by filename. Every row links to the original photo it was made from whenever that original is in the dataset. Load it No Hugging Face account or token is needed: from datasets import load_dataset ds = load_dataset("buildchange/segmentation_training_data", split="train") row = ds[0] row["image"]… See the full description on the dataset page: https://huggingface.co/datasets/buildchange/segmentation_training_data.
Build Change augmented segmentation dataset
Paired images and segmentation masks, including augmented (rotated, flipped, transformed) copies of each original photo, paired by filename. Every row links to the original photo it was made from whenever that original is in the dataset.
Load it
No Hugging Face account or token is needed:
from datasets import load_dataset
ds = load_dataset("buildchange/segmentation_training_data", split="train")
row = ds[0]
row["image"], row["mask"]Columns
When base_filename is set it is an original of the same source_id in the dataset. To get the original photo for a row:
index_of = {name: i for i, name in enumerate(ds["filename"])}
if row["base_filename"]:
original = ds[index_of[row["base_filename"]]]Splitting into train / test: group by source_id
Edited copies are the same photo as their original. Splitting rows randomly lets a flipped copy land in the test set while its original is in the training set, which inflates test scores. Split by source_id so every copy of a photo lands on the same side.
Provenance
Built by data_upload/upload.py from Google Drive image and mask folders. The first folders are in data/; every later batch is in data/<batch>/. Each upload's pairing report is in pairing_report.json or reports/<batch>.json.
