Team Ai
Datasetpublic

AdaptiveChunking/hnet-chunking-probes

H-Net chunker boundary probes Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical FLORES+ probe at every checkpoint. This is the raw material for studying when a learned segmentation stabilises. Layout <run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte offsets can be aligned to external gold data). 40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000. The early… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-probes.

sourceHugging Faceapache-2.0updated 24d agoView on Hugging Face
0likes4.4kdownloads
Dataset Card

H-Net chunker boundary probes

Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical FLORES+ probe at every checkpoint. This is the raw material for studying when a learned segmentation stabilises.

Layout

<run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte offsets can be aligned to external gold data).

  • —40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000. The early log-spacing is deliberate — uniform 200-step spacing misses the sharp early transitions, and it cannot be recovered after a run.
  • —12 languages: de, el, en, fi, hi, my, ru, ta, th, tr, vi, zh
  • —481 files per run (jitter_* runs have 157; they were logged at a coarser cadence).

npz keys

keyshapemeaning
mask(L,) uint81 = a unit boundary falls before this byte
score(L,) float32the router's continuous boundary probability
sent_offsets(n_sent+1,) int64byte offset of each probe sentence
bpbscalarper-language bits-per-byte at this checkpoint
stepscalartraining step
b_pad, p_pad, pad_mask(128, 315)per-sentence padded views

Known issue: mask and mask_level1 (and score/score_level1) are byte-for-byte identical — the two-level logging wrote the same level twice, so per-level word-vs-morph typing cannot be done from these dumps. For Thai, len(mask) exceeds the probe byte length by 21 bytes. Both are logging bugs, recorded here rather than silently patched.

Caveats worth knowing before you use these

  • —Boundaries are phase-locked to a seed-arbitrary sub-character byte position. Chunk starts land on a UTF-8 character boundary 29–31% of the time against a 33% chance rate (verified independently against the span dumps, not just the mask).
  • —20 of 36 language-seed pairs have not converged by step 6000, so a step-6000 analysis is measuring an unconverged segmentation for the low-resource half.
  • —Probe text decoded with errors="ignore" no longer byte-aligns to the mask for truncated sentences (up to 35/128 for my); drop those rather than assume alignment.

Companion models and results.

AdaptiveChunking/hnet-chunking-probes · Team Ai