ssg1/places-water-binary
Places — Water or No Water 34 original photographs labelled by whether a body of water is visible, resized to 224x224. Homework 1. Property Value Splits train (391), validation (5), test (6) Resolution 224x224 RGB Target label — 1 water, 0 no_water Balance (all originals) 17 water / 17 no_water Purpose Binary image classification: is there a body of water in this scene? Composition Column Type Description image image… See the full description on the dataset page: https://huggingface.co/datasets/ssg1/places-water-binary.
Places — Water or No Water
34 original photographs labelled by whether a body of water is visible, resized to 224x224. Homework 1.
Purpose
Binary image classification: is there a body of water in this scene?
Composition
Collection
All photographs are my own, taken on an iPhone and selected from my camera roll. Nothing was downloaded or generated.
A photograph is water when a body of water is visibly identifiable: river, lake, ocean, waterfall, pond, canal, or fountain. Labels were assigned by hand by filing each photograph into a water/ or no_water/ folder.
Selection kept water independent of setting. Both classes contain built-up scenes and open landscape — skylines over water and a night fountain on the water side, a mountain valley and tree-lined paths on the other. Otherwise a model could score well by classifying "nature vs. city" without ever looking for water.
Preprocessing
Largest square crop, then LANCZOS resize to 224x224 RGB. Cropping rather than squashing keeps the scene undistorted. Padding was considered but rejected: the discarded margin is sky or foreground in every photograph, so no evidence behind a label is lost, and letterbox bars would only shrink the subject.
One exception: IMG_0498 frames a skyline high with the water along the bottom edge, so a centre crop would remove the water that defines its label. It uses a bottom-anchored crop, discarding only empty sky. crop_mode records which rule each row used (33 centre, 1 bottom).
Augmentation
16 variants per training image (368 total), cycling through four transforms with seed 24679. Validation and test photographs are never used as parents:
Each transform changes how a scene was captured, not what is in it, so labels are inherited from the parent and never recomputed. Random crops, heavy blur and full desaturation were rejected — each can remove or obscure the water that defines the label.
Splits and intended use
Photographs were split about 70/15/15, stratified on the label, before augmentation, so no validation or test image has a descendant in train. Use the splits as given.
from datasets import load_dataset
ds = load_dataset("ssg1/places-water-binary")
train, validation, test = ds["train"], ds["validation"], ds["test"]Images are already 224x224, so they feed pretrained backbones without further resizing. Normalisation is left to the user.
Limitations
- 34 original images is very small, so
validationholds 5 images andtestholds 6. A single misclassification moves test accuracy by roughly 17 points — treat any single score as a rough indication rather than a reliable estimate. - Augmentation adds variation, not new scenes, so training-set performance is optimistic.
- Photographer bias — one person, one phone, mostly daylight. Night scenes and bad weather are thinly covered.
- Water scale varies widely — an ocean filling half the frame versus a small fountain. Small-water cases are underrepresented.
- Water type is not labelled. Ocean, river, pond and fountain all collapse to
1. - Built for coursework, not for safety or monitoring use.
Ethical notes
No photograph is a portrait and none has a face as its subject. Some street and park scenes contain distant pedestrians in public places, incidental to the scene and not identifiable at 224x224. No readable licence plates, house numbers or screens appear.
All images are my own work, so there are no third-party copyright questions.
License
MIT.
AI usage disclosure
- Photographs — taken and selected by the author. No images were generated or edited by an AI model.
- Preprocessing and augmentation — deterministic Pillow and NumPy operations, reproducible from seed 24679. No generative model was used.
- Notebook and card — Claude helped structure the notebook and draft documentation. Subject, selection, labelling, the crop exception, and the choice of transforms were decided and verified by the author.
