Team Ai
Datasetpublic

danielrosehill/jerusalem-poster-detection

Jerusalem street poster sightings — detection dataset Geotagged street photographs of a messianic postering campaign in central Jerusalem, with bounding boxes around every campaign poster in frame. Each image carries the GPS coordinates it was shot at, so the dataset supports both detection and spatial analysis. Each class is a specific poster design, not a generic "poster" category. The intent is a recogniser that answers which known artwork is on the wall and where its bounds… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/jerusalem-poster-detection.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes188downloads
Dataset Card

Jerusalem street poster sightings — detection dataset

Geotagged street photographs of a messianic postering campaign in central Jerusalem, with bounding boxes around every campaign poster in frame. Each image carries the GPS coordinates it was shot at, so the dataset supports both detection and spatial analysis.

Each class is a specific poster design, not a generic "poster" category. The intent is a recogniser that answers which known artwork is on the wall and where its bounds are, so the class list grows as new designs are surveyed rather than being redefined.

chabad-rebbe-poster-1 — Chabad Rebbe Poster 1

Portrait of the Lubavitcher Rebbe, Menachem Mendel Schneerson, in a black fedora against a yellow ground, captioned in Hebrew יחי אדוננו מורנו ורבינו מלך המשיח לעולם ועד -- commonly shortened to יחי המלך המשיח, 'long live King Messiah'. Some reproductions carry an additional line, ברכה והצלחה ('blessing and success'), above the portrait.

  • —Class index 0 · 26 instance(s) in this release
  • —Physical forms: wheatpaste, sticker
  • —First recorded: 2026-08-31
  • —Reproduced at two very different scales: large wheatpasted sheets on construction hoardings, and small yellow stickers on lampposts and utility boxes. Same artwork, so one identifier.
  • —Observed intact, over-sprayed with black paint, torn, and sun-bleached to near-illegibility. All states are labelled as this design.

At a glance

Images13
Annotated instances26
Classes1 (chabad-rebbe-poster-1)
Splitstrain
Image size1920x2560 (portrait)
Captured2026-08-31
Area31.7822-31.7841 N, 35.2159-35.2195 E
Source repository<https://github.com/danielrosehill/Jerusalem-Graffiti>

Class identifiers

Class indices come from annotations/designs.json in the source repository and are append-only. A new design gets the next index; existing indices never move. Reordering them would silently change the meaning of every published label file and every model trained on an earlier release, so a released index is permanent even if the design is later withdrawn.

Designs seen but not yet annotated are tracked separately in that file and are deliberately not classes — a design becomes a class only when it has boxes.

Format

Hugging Face imagefolder layout: data/<split>/*.jpg alongside a metadata.jsonl keyed on file_name.

python
from datasets import load_dataset
ds = load_dataset("imagefolder", data_dir="data")

Boxes are COCO-style `[x, y, width, height]` in absolute pixels, top-left origin — not normalised, and not [x1, y1, x2, y2]. Getting this wrong is the usual cause of a model that trains without error and detects nothing.

json
{"file_name": "...jpg", "objects": {"id": [0], "category": [0],
  "bbox": [[842.0, 898.0, 114.0, 239.0]], "area": [27246.0]}}

An Ultralytics YOLO tree (normalised cx cy w h, with data.yaml) is emitted alongside when the dataset is built with --yolo.

Per-image metadata

Beyond the boxes, every row carries the survey fields: latitude, longitude, altitude_m, captured_date, captured_time, location_id, form (wheatpaste / sticker / mixed), mounting (lamppost, hoarding, wall, ...), condition (intact, sprayed, torn, faded), status, duplicate_of and a free-text notes field.

form, mounting and condition describe the image as a whole, not individual boxes. A two-class wheatpaste/sticker variant can be derived from form without re-annotating, except on rows marked mixed.

Splitting

Split on `location_id`, not on rows. Several frames show the same lamppost from a few paces apart — duplicate_of marks the clearest cases — so a random per-image split puts near-identical photographs in both train and validation and reports a score that is largely memorisation. This dataset ships as a single train split for that reason; do your own grouped split or grouped cross-validation.

Annotation protocol

Boxes were drawn by hand in labelme over the full-resolution originals, one box per visible campaign poster.

  • —Partly occluded posters are boxed at their visible extent.
  • —Posters too blurred or distant to identify with confidence are not boxed; notes records several such judgements explicitly.
  • —Non-campaign material that looks similar — memorial notices, election flyers, blue-and-white "baruch haba" stickers of a different design, plain utility stickers — is deliberately left unboxed rather than labelled. Those are the confusions a detector will actually make, so they matter as background.
  • —0 image(s) contain no box. These are verified negatives, looked at and found to hold nothing, not unannotated files.

Instances are small relative to the frame: the smallest box is about 65 px on a side and the median about 145 px, in 1920x2560 images. Detectors trained at 640 px will lose the smallest stickers entirely; train at higher resolution or on tiles.

Limitations

Read these before training anything you intend to trust.

  • —It is small. 13 images and 26 instances, from a single walk. The working threshold for a genuine YOLO fine-tune is roughly 50 images and 150+ instances; this is below it. Treat it as a seed set, a benchmark for few-shot and open-vocabulary detectors, or a starting point to extend.
  • —One photographer, one route, one time of day. Late-afternoon light, one camera, one stretch of one street. No night, rain, or winter frames.
  • —Correlated backgrounds. Jaffa Road street furniture recurs constantly, so a model can learn "lamppost" as a proxy for "poster".
  • —Instances are small and often high on poles, frequently against blown-out sky, and several are damaged — over-sprayed, torn, or sun-bleached to near-illegibility. That damage is the interesting part of the survey and it is deliberately in-distribution here.
  • —Not a census. These are the sightings one person recorded, not every poster in the area, so counts carry no denominator.

Provenance and intent

The photographs document postering in shared public space; the survey exists because a municipal helpline had not acted on reports of it. Images are published as shot — faces and number plates are not redacted, and passers-by appear incidentally, as they do in ordinary street photography. If you appear in a frame and want it removed, open an issue on the source repository.

The dataset records where campaign material was found. It is not evidence about any individual, and nothing in it identifies who put the posters up. Do not use it to make claims about people who happen to be visible in the frames.

Citation

bibtex
@misc{jerusalem_poster_survey,
  title  = {Jerusalem messianic poster sightings: a geotagged detection dataset},
  author = {Rosehill, Daniel},
  year   = {2026},
  url    = {https://github.com/danielrosehill/Jerusalem-Graffiti}
}

Licence

Images and annotations: CC BY 4.0.