Team Ai
Datasetpublic

langminer/watermark-benchmark-v18

Multi-Class Watermark & Camera Stamp Dataset (Round 18) This repository contains the complete dataset, augmentation assets, and real-world evaluation benchmarks used to train the Champion 3-Class Watermark Classifier (wm_3class_v18_scratch.pt). The dataset addresses a critical problem in media ingestion: automatically excluding produced media, broadcast stills, and stock photography without falsely excluding authentic personal photographs (0.00% false alarms on personal photos… See the full description on the dataset page: https://huggingface.co/datasets/langminer/watermark-benchmark-v18.

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes258downloads
Dataset Card

Multi-Class Watermark & Camera Stamp Dataset (Round 18)

This repository contains the complete dataset, augmentation assets, and real-world evaluation benchmarks used to train the Champion 3-Class Watermark Classifier (wm_3class_v18_scratch.pt).

The dataset addresses a critical problem in media ingestion: automatically excluding produced media, broadcast stills, and stock photography without falsely excluding authentic personal photographs (0.00% false alarms on personal photos and challenging hard negatives).


1. Dataset Overview

  • —Total Samples: 65,733 images across 21,911 aligned triplet pairs.
  • —Train Split: 56,280 images (18,760 triplets)
  • —Validation Split: 9,453 images (3,151 triplets)
  • —Classes (3-way):
  • —0: clean — Authentic personal photographs, public events, street scenes, and challenging hard negatives (foliage, glass reflections, barcodes).
  • —1: publisher — Publisher watermarks, TV channel bugs (BBC News, PBS NewsHour, Fox News), stock photography previews (Dreamstime, Getty, Shutterstock), and agency credits.
  • —2: camera — Authentic smartphone camera timestamps and model watermarks (Huawei, Oppo, Samsung, Xiaomi, Vivo, Tecno, etc.).
  • —Aligned Triplet Structure: Every sample belongs to an aligned triplet sharing the exact same base canvas image, ensuring the classifier isolates the watermark signal rather than background scene semantics.

2. Directory Structure

.
├── README.md                      # Hugging Face Dataset Card
├── data/                          # Sharded Apache Parquet files with embedded images
│   ├── train-00000-of-00010.parquet
│   ...
│   └── validation-00001-of-00002.parquet
├── augmentation/                  # Raw assets used to synthesize the dataset
│   ├── marks/                     # 2,796 transparent PNG overlay marks
│   │   └── manifest.json          # Metadata for all vector SVGs and procedural marks
│   ├── bases/                     # 6,528 clean base canvas photographs
│   ├── bases_manifest.json        # Base image mappings and origin categories
│   └── scripts/                   # Augmentation scripts (generate_pairs.py, build_marks.py)
├── benchmarks/                    # Real-world evaluation benchmarks
│   ├── news_broadcast_stills/     # 139 authentic TV news broadcast frames
│   ├── hard_negatives/            # 880 Wikimedia challenging photos (glass, foliage, barcodes)
│   ├── camera_watermarks/         # 131 harvested handset camera watermark test photos
│   └── glm_verified_testset/      # 259 GLM-verified gold standard benchmark images & labels
└── upload_to_hf.py                # 1-click script to upload to Hugging Face Hub

3. How to Use with Hugging Face datasets

python
from datasets import load_dataset

# Load dataset (streaming or local download)
ds = load_dataset("langminer/watermark-benchmark-v18")

print(ds)
# DatasetDict({
#     train: Dataset({features: ['image', 'label', 'label_name', 'pair_id', 'split', 'style', 'mark', 'kind', 'opacity', 'tint', 'bbox', 'base_image'], num_rows: 56280}),
#     validation: Dataset({features: ['image', 'label', 'label_name', 'pair_id', 'split', 'style', 'mark', 'kind', 'opacity', 'tint', 'bbox', 'base_image'], num_rows: 9453})
# })

# Access a sample
sample = ds["train"][0]
image = sample["image"]       # PIL Image object
label = sample["label_name"]  # 'clean', 'publisher', or 'camera'
bbox = sample["bbox"]         # [ymin, xmin, ymax, xmax] normalized
image.show()

4. Benchmark Results on Champion Model (v18)

Benchmark SetRound 16 (`v16`)Round 17 (`v17`)**Round 18 (`v18`)**
Dev Selection Metric ($F_{0.5}$)$0.9723$$0.9705$`0.9742`
Dev Precision / Recall$97.1\% / 93.5\%$$97.6\% / 94.7\%$`98.2% / 94.5%`
GLM-Verified Test MCC (259 images)$0.9241$$0.9380$`0.9541`
GLM-Verified Test AUC$0.9972$$0.9972$`0.9951`
News Broadcast Bugs Recall$34.4\%$$49.0\%$`72.9%` (78.1% FP32)
Real Personal Album FA (6,365 photos)$0.00\%$$0.00\%$`0.00%` @ Cut +5.0
Challenging Hard Negatives (880 photos)$0.00\%$$0.00\%$`0.00%` @ Cut +5.0

5. License & Citation

The synthetic dataset, annotations, and marks are released under the Apache 2.0 License. Base photographs originate from Creative Commons sources (Wikimedia, Flickr CC-BY, and public domain collections).

langminer/watermark-benchmark-v18 · Team Ai