langminer/watermark-benchmark-v18
Multi-Class Watermark & Camera Stamp Dataset (Round 18) This repository contains the complete dataset, augmentation assets, and real-world evaluation benchmarks used to train the Champion 3-Class Watermark Classifier (wm_3class_v18_scratch.pt). The dataset addresses a critical problem in media ingestion: automatically excluding produced media, broadcast stills, and stock photography without falsely excluding authentic personal photographs (0.00% false alarms on personal photos… See the full description on the dataset page: https://huggingface.co/datasets/langminer/watermark-benchmark-v18.
Multi-Class Watermark & Camera Stamp Dataset (Round 18)
This repository contains the complete dataset, augmentation assets, and real-world evaluation benchmarks used to train the Champion 3-Class Watermark Classifier (wm_3class_v18_scratch.pt).
The dataset addresses a critical problem in media ingestion: automatically excluding produced media, broadcast stills, and stock photography without falsely excluding authentic personal photographs (0.00% false alarms on personal photos and challenging hard negatives).
1. Dataset Overview
- Total Samples: 65,733 images across 21,911 aligned triplet pairs.
- Train Split: 56,280 images (18,760 triplets)
- Validation Split: 9,453 images (3,151 triplets)
- Classes (3-way):
0: clean— Authentic personal photographs, public events, street scenes, and challenging hard negatives (foliage, glass reflections, barcodes).1: publisher— Publisher watermarks, TV channel bugs (BBC News, PBS NewsHour, Fox News), stock photography previews (Dreamstime, Getty, Shutterstock), and agency credits.2: camera— Authentic smartphone camera timestamps and model watermarks (Huawei, Oppo, Samsung, Xiaomi, Vivo, Tecno, etc.).- Aligned Triplet Structure: Every sample belongs to an aligned triplet sharing the exact same base canvas image, ensuring the classifier isolates the watermark signal rather than background scene semantics.
2. Directory Structure
.
├── README.md # Hugging Face Dataset Card
├── data/ # Sharded Apache Parquet files with embedded images
│ ├── train-00000-of-00010.parquet
│ ...
│ └── validation-00001-of-00002.parquet
├── augmentation/ # Raw assets used to synthesize the dataset
│ ├── marks/ # 2,796 transparent PNG overlay marks
│ │ └── manifest.json # Metadata for all vector SVGs and procedural marks
│ ├── bases/ # 6,528 clean base canvas photographs
│ ├── bases_manifest.json # Base image mappings and origin categories
│ └── scripts/ # Augmentation scripts (generate_pairs.py, build_marks.py)
├── benchmarks/ # Real-world evaluation benchmarks
│ ├── news_broadcast_stills/ # 139 authentic TV news broadcast frames
│ ├── hard_negatives/ # 880 Wikimedia challenging photos (glass, foliage, barcodes)
│ ├── camera_watermarks/ # 131 harvested handset camera watermark test photos
│ └── glm_verified_testset/ # 259 GLM-verified gold standard benchmark images & labels
└── upload_to_hf.py # 1-click script to upload to Hugging Face Hub3. How to Use with Hugging Face datasets
from datasets import load_dataset
# Load dataset (streaming or local download)
ds = load_dataset("langminer/watermark-benchmark-v18")
print(ds)
# DatasetDict({
# train: Dataset({features: ['image', 'label', 'label_name', 'pair_id', 'split', 'style', 'mark', 'kind', 'opacity', 'tint', 'bbox', 'base_image'], num_rows: 56280}),
# validation: Dataset({features: ['image', 'label', 'label_name', 'pair_id', 'split', 'style', 'mark', 'kind', 'opacity', 'tint', 'bbox', 'base_image'], num_rows: 9453})
# })
# Access a sample
sample = ds["train"][0]
image = sample["image"] # PIL Image object
label = sample["label_name"] # 'clean', 'publisher', or 'camera'
bbox = sample["bbox"] # [ymin, xmin, ymax, xmax] normalized
image.show()4. Benchmark Results on Champion Model (v18)
5. License & Citation
The synthetic dataset, annotations, and marks are released under the Apache 2.0 License. Base photographs originate from Creative Commons sources (Wikimedia, Flickr CC-BY, and public domain collections).
