AbstractPhil/beatrix-captured-interactive-inferences
Beatrix captured interactive inferences Byte-by-byte internals of real inference runs on mini-beatrix-2.5s, the 237.1M full-splat byte model (model). Nothing here is simulated or sampled from a proxy: each capture is one greedy generation, re-run through a single instrumented forward pass that walks the blocks by hand and records what every layer did at every byte. Each prompt is captured twice over the same byte sequence — once on the bare core and once with the library's top… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/beatrix-captured-interactive-inferences.
Beatrix captured interactive inferences
Byte-by-byte internals of real inference runs on mini-beatrix-2.5s, the 237.1M full-splat byte model (model). Nothing here is simulated or sampled from a proxy: each capture is one greedy generation, re-run through a single instrumented forward pass that walks the blocks by hand and records what every layer did at every byte.
Each prompt is captured twice over the same byte sequence — once on the bare core and once with the library's top arm (rules) mounted — so the arm's contribution is a difference at every byte and layer, not an inference from two separate runs.
Open it
[Beatrix Byte Scope](https://claude.ai/artifact/YH3BfrBWiPM5QfsgDuuKuJ) — the interactive viewer for this data, and the official public copy. A byte strip you click, slide or play through; a depth-by-byte map over all twenty blocks; the effective attention row and the layer's full causal map; the address read and the per-byte write across all four books at once or one at a time; the board's width in use over the whole run; the anchored experts; the head; the codebook frames. A three-way switch recomputes every panel for the bare core, the core with the arm, or the difference.
The page's source is viewer/scope.html here, beside the data it reads. To run it from a clone, serve the repo root and open the file through that server — a file:// page cannot read its own folder:
git clone https://huggingface.co/datasets/AbstractPhil/beatrix-captured-interactive-inferences
cd beatrix-captured-interactive-inferences && python -m http.server 8000
# then open http://localhost:8000/viewer/scope.htmlCaptures
Layout
manifest.json every capture's bytes, replies and metadata,
plus per-layer codebook health
<capture>/<core|arm>/
summary.json per-byte scalars, per-layer scalars, shapes,
means and scales
read.u8 (layers, books, bytes, anchors) the signed read
write.u8 same shape: what THIS byte writes
massn.u8 same shape: the board's shape, normalized
mass.u8 same shape: the raw cumulative load
attn.u8 (layers, bytes, bytes) effective attention
head.u8 (bytes, 256) the head's signed address
arm.u8 (layers, bytes, 64) the arm's read (arm only)
viewer/scope.html the interactive page
capture.py the script that produced all of itReading the arrays
All binaries are raw uint8, C order, no header. `read`, `write`, `massn`, `head` and `arm` are stored as a mean plus a deviation —
value[l][c][i][k] = mean[l][c][k] + (u8 - 128)/127 * scale[l][c]with mean and scale in summary.json. mass alone keeps the plain absolute form, value = u8/255 * scales.mass.
import numpy as np, json
s = json.load(open("rule-chain/arm/summary.json"))
L, C, n, K = s["shapes"]["read"]
q = np.fromfile("rule-chain/arm/read.u8", dtype=np.uint8).reshape(L, C, n, K)
dev = (q.astype(np.float32) - 128) / 127 * np.array(s["scales"]["read"])[:, :, None, None]
val = np.array(s["mean"]["read"])[:, :, None, :] + devWhy the split, and why it matters. Every anchor picture is dominated by a large constant component: the cosine similarity between neighbouring bytes is 1.000, and the part that varies byte to byte is about 4% of the whole. An earlier version of this dataset quantized the absolute value against one global scale, which put that varying part below a single quantization step — the arrays looked right and carried almost no per-byte information. Splitting the constant off restores it. The scale is the 99.5th percentile of |deviation|, not the maximum, because the first two bytes of a sequence deviate ~25x more than every later byte (the board is nearly empty there) and against a max-based scale those two bytes crushed all the others into 25 of the 256 levels. summary.json reports the clipped fraction. For a per-byte reading, use the deviation directly; add the mean back only when you want the absolute value.
attn.u8 is row-normalized to its own maximum, so a row reads as a relative profile over earlier bytes rather than an absolute weight.
summary.json also carries, per byte and layer: the residual norm in and out, the attention contribution, the bank's trunk and dispatched-expert contributions, the three experts' signed dispatch weights, the splat agreement mass per book, the board's effective width in anchors (used), and (arm runs) the patch norm and gate. Per byte at the head: the top eight predicted bytes with probabilities, the entropy in bits, and the probability given to the byte actually taken.
What the arrays are
The splat attention has no softmax over positions. Each layer reads a fixed-width addressed blackboard: a query's oriented halves are read against the accumulated mass, and the result is divided by the scalar agreement mass. read is the query's signed per-anchor weight, write is what this byte puts on the board, mass is the running sum of those writes, and attn is the exact effective byte-to-byte weight implied by the same bilinear forms the scan sums, so it is derived rather than approximated:
att[i,j] = sum_books ( qp_i . kp_j + qn_i . kn_j ), j <= i
divided by den_iThe board is a running sum, so plotting it directly plots position in the sequence and little else. The two readings of it that move are the per-byte write and the board's effective width (used), which in block 10 of the rule capture falls from 36 anchors at the first byte to under 2 by byte 40 and stays there.
The read is signed, so a negative weight is inhibition — a first-class result of the closed form, not an absence. The viewer draws it in teal against amber.
Measured fp32, greedy, on one RTX 4090.
