AdaptiveChunking/hnet-chunking-probes
H-Net chunker boundary probes Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical FLORES+ probe at every checkpoint. This is the raw material for studying when a learned segmentation stabilises. Layout <run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte offsets can be aligned to external gold data). 40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000. The early… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-probes.
H-Net chunker boundary probes
Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical FLORES+ probe at every checkpoint. This is the raw material for studying when a learned segmentation stabilises.
Layout
<run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte offsets can be aligned to external gold data).
- 40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000. The early log-spacing is deliberate — uniform 200-step spacing misses the sharp early transitions, and it cannot be recovered after a run.
- 12 languages: de, el, en, fi, hi, my, ru, ta, th, tr, vi, zh
- 481 files per run (
jitter_*runs have 157; they were logged at a coarser cadence).
npz keys
Known issue: mask and mask_level1 (and score/score_level1) are byte-for-byte identical — the two-level logging wrote the same level twice, so per-level word-vs-morph typing cannot be done from these dumps. For Thai, len(mask) exceeds the probe byte length by 21 bytes. Both are logging bugs, recorded here rather than silently patched.
Caveats worth knowing before you use these
- Boundaries are phase-locked to a seed-arbitrary sub-character byte position. Chunk starts land on a UTF-8 character boundary 29–31% of the time against a 33% chance rate (verified independently against the span dumps, not just the mask).
- 20 of 36 language-seed pairs have not converged by step 6000, so a step-6000 analysis is measuring an unconverged segmentation for the low-resource half.
- Probe text decoded with
errors="ignore"no longer byte-aligns to the mask for truncated sentences (up to 35/128 formy); drop those rather than assume alignment.
