einarolafsson/spacr-example-annotate
spaCR — Annotate and Classify example data Example input for the Annotate and Classify modules of spaCR. It is the output of a Measure run, so both modules can be exercised without segmenting or measuring anything first. What is here Path What it is data/ 2,341 single-cell PNG crops, foldered by phenotype measurements.db The measurements, plus png_list and the annotation tables measurements/active_learning/ The model card from the first annotation… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-example-annotate.
spaCR — Annotate and Classify example data
Example input for the Annotate and Classify modules of spaCR. It is the output of a Measure run, so both modules can be exercised without segmenting or measuring anything first.
What is here
Labels: infected or not, by rule
png_list.infected holds a label for every one of the 2,341 crops, derived from the measurements rather than by hand:
1,173 infected against 1,168 uninfected — 50.1%, which is about as balanced as a binary training set gets without being resampled.
How it was derived. The pathogen table names the parent cell of every detected parasite, so a cell is infected exactly when at least one pathogen row points at it. That is a rule anyone can re-run and check, which a hand pass over a subset is not.
Independently cross-checked. All 1,173 also have cell_pathogen_overlap_fraction > 0 — a column Measure computes by a different route entirely. The two agree on every cell.
annotate is deliberately empty, so the module opens on a clean column and the rule-based labels stay as a reference rather than as something to overwrite.
Paths are RELATIVE, deliberately
Every path in measurements.db and in both settings files is relative to the dataset root:
data/single_nucleus/uninfected/plate1_E01/cell_png/plate1_E01_19_1_12.pngA measurements database normally stores absolute paths, which name the machine that made it and resolve nowhere else. spaCR's downloader rewrites these to absolute on arrival, so the database works wherever it is unpacked. If you unpack it by hand, do the same, or point spaCR at the folder and let it.
<dataset> in the settings files is a placeholder for the unpack location and is substituted the same way.
Plate layout
Four wells (E01, E02, L01, L02) from the plate published as `einarolafsson/spacr-example-measure`, which holds the merged arrays these crops were cut from.
How spaCR downloads it
Everything here is also published as a single uncompressed `spacr-example-annotate.tar`, and that is what spaCR fetches: one request instead of 2,365, a progress figure that means something, and a download that stops when you press Cancel. It is unpacked with tar's data filter, which refuses any member that would write outside the destination folder.
The individual files are kept beside it so the set can be browsed and previewed on this page. They are the same bytes; either is fine to use.
Provenance
Produced by spaCR's Mask module and then its Measure module, from spaCR's own example images, with cpsam for cells and nuclei and a Toxoplasma-specific CPSAM checkpoint for pathogens. One field of fifty-two (plate1_E02_20_1) failed to measure — a nucleus label spanning two cells — so the crops come from 51 fields. Absolute paths have been rewritten to relative throughout.
