Team Ai
Datasetpublic

buildchange/segmentation_training_data

Build Change augmented segmentation dataset Paired images and segmentation masks, including augmented (rotated, flipped, transformed) copies of each original photo, paired by filename. Every row links to the original photo it was made from whenever that original is in the dataset. Load it No Hugging Face account or token is needed: from datasets import load_dataset ds = load_dataset("buildchange/segmentation_training_data", split="train") row = ds[0] row["image"]… See the full description on the dataset page: https://huggingface.co/datasets/buildchange/segmentation_training_data.

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes291downloads
Dataset Card

Build Change augmented segmentation dataset

Paired images and segmentation masks, including augmented (rotated, flipped, transformed) copies of each original photo, paired by filename. Every row links to the original photo it was made from whenever that original is in the dataset.

Load it

No Hugging Face account or token is needed:

python
from datasets import load_dataset

ds = load_dataset("buildchange/segmentation_training_data", split="train")
row = ds[0]
row["image"], row["mask"]

Columns

columnmeaning
imagethe photo
maskits segmentation mask; a few are a zoomed-out copy at a smaller size, so resize the mask to the image before training
filenamesource filename, e.g. b13_croped1_rotated_350.png
source_idthe real-world photo this row came from, e.g. b13_croped1
base_filenamefilename of the unedited original, e.g. b13_croped1.png (an original points at itself; empty when the original is not in the dataset)
is_originalTrue for the unedited photo, False for an edited copy
augmentationthe edit applied, e.g. vflip, rotated_350; empty for originals

When base_filename is set it is an original of the same source_id in the dataset. To get the original photo for a row:

python
index_of = {name: i for i, name in enumerate(ds["filename"])}
if row["base_filename"]:
    original = ds[index_of[row["base_filename"]]]

Splitting into train / test: group by source_id

Edited copies are the same photo as their original. Splitting rows randomly lets a flipped copy land in the test set while its original is in the training set, which inflates test scores. Split by source_id so every copy of a photo lands on the same side.

Provenance

Built by data_upload/upload.py from Google Drive image and mask folders. The first folders are in data/; every later batch is in data/<batch>/. Each upload's pairing report is in pairing_report.json or reports/<batch>.json.