Team Ai
Modelpublic

AbstractPhil/mini-beatrix-3

sourceHugging Facecc-by-nc-sa-4.0updated 2d agoView on Hugging Face
0likes1.4kdownloads
Model Card

mini-beatrix-3

The 32-block rung of the mini-beatrix ladder, with its stage arms. 376M parameters, byte-level (vocab 256), with a multi-constellation CausalSplatHUB (signed-address linear attention over learned codebook blackboards) and an anchored expert bank in every one of its 32 blocks. No softmax-over-positions attention anywhere. Trained on 64.4B bytes (245,674 steps, about 295 hours on two RTX 5090s) through a staged curriculum, completed 2026-10-05.

The package ships both states in one repository: the bare model (arms off) and the model with its stage arms mounted (arms on). An arm is a small detachable adapter trained on this exact frozen core. The arms come as groups: arms that were trained switched on together and are mounted together. There are two, the four arms of stages 1 to 4 and the eight arms of stages 1 to 8 (those four, held fixed, with four more fitted over them), and an experimental third: the eight, held fixed, with a ninth arm for image captions fitted over them, a prototype that is listed and mountable like the others but is not the default. Each stage also has a solo arm, trained alone, for use one at a time. Arms are guests on the core, never a change to it: the core's weights are the same files either way, and detaching restores the bare model bit for bit.

python
import torch
from transformers import AutoModelForCausalLM

m = AutoModelForCausalLM.from_pretrained(
        "AbstractPhil/mini-beatrix-3", trust_remote_code=True).eval()

m.say("Who are you?")            # arms off: the bare core, its own chat frame

m.mount_arm("stages-1-8")         # arms on: the stage arms, all on together
m.detach_arm()                   # the bare core again, verified bit-exact

Arms on from the first call:

python
m = AutoModelForCausalLM.from_pretrained(
        "AbstractPhil/mini-beatrix-3", trust_remote_code=True,
        default_arm="stages-1-8").eval()

The model reads raw UTF-8 bytes: input_ids are byte values 0–255. There is no tokenizer to download.

Architecture

  • —d_model 1024 · 32 layers · 16 heads · ctx 4096 · byte-trigram embedding (raw UTF-8 bytes; input ids are byte values 0–255)
  • —Hubs (all 32 blocks): 4 constellations × 64 anchors @ D=128 per block, read by an exact chunked scan (chunk 256) and budget-composed (numerators and agreement masses sum before one divide: reconstructive, never comparative, with no argmax and no top-k). Constant-size prefix state: each layer encodes the sequence onto a fixed-width addressed blackboard rather than caching it.
  • —Anchored banks: 3 full-width experts per block (ff 1024), signed dispatch, expert outputs born at zero.
  • —Dual head: linear readout + a signed aleph read (256 anchors @ 256), fitted to the first batch at step 0 so that it works from birth (8.26 → 5.45 bpb on that batch).
  • —Held-out cost of switching a mechanism off (the toggle ledger, in bpb), at the last boundary before the end and at the end: hubs +5.86 and +6.41 · head +1.94 and +1.90 · banks +3.40 at the former. The banks' reading at the end (+6.13) is nearly twice every earlier one and is being repeated before it is relied on.

The stage arms

The curriculum taught the model nine kinds of text, one stage after another, and the finished model moved on from each stage's text as the later stages came. The stage arms bring that text back without touching the core. They were fitted on the finished, frozen core as always-on groups: the arms of stages 1 and 2 were trained first and then held fixed, the arms of stages 3 and 4 were attached over them one after the other and trained in their presence, and then, with all four held fixed and on, the arms of stages 5, 6, 7 and 8 were attached over them the same way, one after the other, each for 4,000 steps and each kept quiet on every other arm's stage text. The eight-arm group's first four arms are the four-arm group unchanged; the four-arm group is also packaged on its own. m.arms() returns the tables below as data, with each row's recipe and full measurements.

The group `stages-1-4` (4 arms, 54.8M parameters; two seeds):

armits stage's textstage text, arms off → the group on (bpb)this arm alone (bpb)stage items in the stage's own form, off → group onstage items in a new form, off → group on
stages-1-4/s1_perspectiveone small event retold from each side: seen from a person, said to them, told about them0.554 → 0.0480.05255% → 96%46% → 48%
stages-1-4/s2_conceptkinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ0.721 → 0.1120.35030% → 98%68% → 78%
stages-1-4/s3_rulesif-then rules over made-up words, followed step by step to what follows0.930 → 0.2240.82416% → 90%7% → 30%
stages-1-4/s4_arithsmall arithmetic worked out in text1.234 → 0.3381.12734% → 61%24% → 24%

With the group on, held-out web text moves by +0.0017 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.419 with the group on.

The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.048, 0.121, 0.222, 0.338 bpb and moved web text by +0.0011.

Each member mounted alone, and the whole group, on every stage's held-out text (bpb):

mountedstage 1 textstage 2 textstage 3 textstage 4 textweb text
nothing (arms off)0.5540.7210.9301.2340.951
only stages-1-4/s1_perspective0.0530.2390.5090.7230.952
only stages-1-4/s2_concept0.5270.3500.6190.8800.952
only stages-1-4/s3_rules0.5540.7240.8241.2340.951
only stages-1-4/s4_arith0.5530.7210.9281.1270.951
the whole group stages-1-40.0480.1120.2240.3380.952

The group `stages-1-8` (8 arms, 109.5M parameters; two seeds):

armits stage's textstage text, arms off → the group on (bpb)this arm alone (bpb)stage items in the stage's own form, off → group onstage items in a new form, off → group on
stages-1-8/s1_perspectiveone small event retold from each side: seen from a person, said to them, told about them0.554 → 0.0490.05255% → 89%46% → 44%
stages-1-8/s2_conceptkinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ0.721 → 0.1110.35030% → 98%68% → 72%
stages-1-8/s3_rulesif-then rules over made-up words, followed step by step to what follows0.930 → 0.2240.82416% → 86%7% → 31%
stages-1-8/s4_arithsmall arithmetic worked out in text1.234 → 0.3281.12734% → 64%24% → 24%
stages-1-8/s5_causalshort cause-and-effect text: what happened and why0.660 → 0.0430.091not measurednot measured
stages-1-8/s6_tryfailan attempt, its failure and the revised attempt0.673 → 0.0550.164not measurednot measured
stages-1-8/s7_mixedthe mixed stage's diet: the earlier stages' text beside ordinary prose0.783 → 0.7200.724not measurednot measured
stages-1-8/s8_registerthe same content in different registers of speech0.491 → 0.0210.068not measurednot measured

With the group on, held-out web text moves by +0.0019 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.422 with the group on.

The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.048, 0.121, 0.222, 0.327, 0.043, 0.054, 0.720, 0.021 bpb and moved web text by +0.0010.

Each member mounted alone, and the whole group, on every stage's held-out text (bpb):

mountedstage 1 textstage 2 textstage 3 textstage 4 textstage 5 textstage 6 textstage 7 textstage 8 textweb text
nothing (arms off)0.5540.7210.9301.2340.6600.6730.7830.4910.951
only stages-1-8/s1_perspective0.0530.2390.5090.7230.4630.4950.7730.3790.952
only stages-1-8/s2_concept0.5270.3500.6190.8800.6590.6730.7830.4690.952
only stages-1-8/s3_rules0.5540.7240.8241.2340.6680.6740.7820.4910.951
only stages-1-8/s4_arith0.5530.7210.9281.1270.6570.6590.7830.4890.951
only stages-1-8/s5_causal0.5360.7100.9171.2250.0910.6660.7820.4840.951
only stages-1-8/s6_tryfail0.5520.7220.9301.2370.6510.1640.7830.4900.951
only stages-1-8/s7_mixed0.4970.7150.8681.1950.5580.5930.7240.4830.950
only stages-1-8/s8_register0.5370.7130.9251.2270.6420.6510.7830.0680.951
the whole group stages-1-80.0490.1100.2240.3280.0430.0550.7200.0210.952

The experimental group `stages-1-8-caption` (9 arms, 123.2M parameters; two seeds; experimental):

Experimental: a prototype for the diffusion line, not a curriculum stage. The ninth arm was fitted for 4,000 steps over the eight (held fixed and on) on image captions, text the core never trained on: five sources, each image one document in a labelled frame (tags, caption, short, attributes, scene, photo, structure, medium, long, rating). With all nine on, held-out caption rows read 0.783 bits per byte against 1.265 with arms off (0.780 on the second seed); the arm alone reads 0.785, so it carries the whole gain itself, and the eight read caption rows like the bare model (within 0.001). With it on, the eight's stage texts lose at most 0.002, web text moves by +0.003 (+0.002 on the second seed; the ninth arm's own share under 0.001) and its largest write on a partner's stage text is +0.002. Its training curve was still falling at 4,000 steps (0.708 to 0.693 over the last 800, where every stage arm had flattened) and its gate opens about twice as wide as any stage arm's (0.45 against 0.11 to 0.28), so a longer arm is an open item. What it does to generated captions, and to diffusion conditioning, is not measured here.

armits stage's textstage text, arms off → the group on (bpb)this arm alone (bpb)stage items in the stage's own form, off → group onstage items in a new form, off → group on
stages-1-8-caption/s1_perspectiveone small event retold from each side: seen from a person, said to them, told about them0.554 → 0.0500.05255% → 86%46% → 44%
stages-1-8-caption/s2_conceptkinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ0.721 → 0.1110.35030% → 96%68% → 70%
stages-1-8-caption/s3_rulesif-then rules over made-up words, followed step by step to what follows0.930 → 0.2250.82416% → 84%7% → 29%
stages-1-8-caption/s4_arithsmall arithmetic worked out in text1.234 → 0.3291.12734% → 65%24% → 25%
stages-1-8-caption/s5_causalshort cause-and-effect text: what happened and why0.660 → 0.0430.091not measurednot measured
stages-1-8-caption/s6_tryfailan attempt, its failure and the revised attempt0.673 → 0.0550.164not measurednot measured
stages-1-8-caption/s7_mixedthe mixed stage's diet: the earlier stages' text beside ordinary prose0.783 → 0.7220.724not measurednot measured
stages-1-8-caption/s8_registerthe same content in different registers of speech0.491 → 0.0220.068not measurednot measured
stages-1-8-caption/s9_captionimage captions in one labelled frame (tags, the source's caption, short, attributes, scene, photo, structure, medium, long, rating: one label per line): photographs with structured captions, CC12M alt text and generated descriptions, COCO captions, danbooru tags1.265 → 0.7830.785not measurednot measured

With the group on, held-out web text moves by +0.0027 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.422 with the group on.

The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.049, 0.122, 0.223, 0.327, 0.043, 0.055, 0.722, 0.021, 0.779 bpb and moved web text by +0.0017.

Each member mounted alone, and the whole group, on every stage's held-out text (bpb):

mountedstage 1 textstage 2 textstage 3 textstage 4 textstage 5 textstage 6 textstage 7 textstage 8 textcaptions textweb text
nothing (arms off)0.5540.7210.9301.2340.6600.6730.7830.4911.2650.951
only stages-1-8-caption/s1_perspective0.0530.2390.5090.7230.4630.4950.7730.3791.2660.952
only stages-1-8-caption/s2_concept0.5270.3500.6190.8800.6590.6730.7830.4691.2660.952
only stages-1-8-caption/s3_rules0.5540.7240.8241.2340.6680.6740.7820.4911.2650.951
only stages-1-8-caption/s4_arith0.5530.7210.9281.1270.6570.6590.7830.4891.2640.951
only stages-1-8-caption/s5_causal0.5360.7100.9171.2250.0910.6660.7820.4841.2640.951
only stages-1-8-caption/s6_tryfail0.5520.7220.9301.2370.6510.1640.7830.4901.2650.951
only stages-1-8-caption/s7_mixed0.4970.7150.8681.1950.5580.5930.7240.4831.2640.950
only stages-1-8-caption/s8_register0.5370.7130.9251.2270.6420.6510.7830.0681.2640.951
only stages-1-8-caption/s9_caption0.5520.7230.9251.2300.6600.6650.7860.5190.7850.952
the whole group stages-1-8-caption0.0500.1110.2240.3290.0430.0550.7220.0220.7830.953

How to read the tables. Stage text is the loss, in bits per byte, on held-out text of the arm's own stage, with arms off and with the whole group on. This arm alone is the same loss with only that one arm mounted. Stage items are short test questions from that stage, scored as the share answered correctly, first in the form the stage text uses and then in a form it does not. The web text figure is the change on held-out ordinary web text (fineweb-edu) with the group on: the arms were trained to leave it alone.

A group is the unit. The eight-arm group is the four-arm group with four more arms fitted over it, and it keeps the four as they were: with all eight on, the text of stages 1 to 4 reads within 0.011 bits per byte of the four-arm group's own reading (stage 4 slightly better, the others the same). The new arms take their stages: with the whole group on, stage 5 text reads 0.043 bits per byte against 0.660 bare, stage 6 0.055 against 0.673, stage 8 0.021 against 0.491, on both seeds. Each of those three carries about two thirds to three quarters of its stage's gain itself and the fixed earlier arms supply the rest (the first arm, a generalist, reads stage 5 and 6 text 0.2 and 0.18 below bare on its own), so again the group works as a whole and is measured that way. The mixed stage (7) is the exception: its text is a blend of the other stages' kinds, the bare model reads it at 0.783, and no arm moves it much (the group reads it at 0.720, its own arm alone at 0.724). The group costs 0.001 to 0.002 bits per byte on ordinary web text, no arm costs more than 0.0006 of that, and no arm writes on another arm's stage text above 0.0012. Nothing an earlier arm learned was lost when a later arm trained over it (the largest loss of an earlier stage's gain at any later close was 2 percent). The four new arms trained for 4,000 steps each, about where a solo arm's curve flattens; the first four kept their shorter training (3,200, 2,400, 800 and 800 steps). For one stage by itself, use its solo arm below.

What the arms do, plainly. A group recovers each of its stages' own kind of text. The gain is in each stage's own forms. Whether it carries over to new forms is shown in each table's last column, where it was measured.

Arms off and arms on

python
m.arm                        # None: arms off
h = m.mount_arm("stages-1-8")       # arms on: the whole eight-arm group
with h.all_off():            # every member masked: the bare core's logits
    m.say("Hello there.")
m.detach_arm(verify=True)    # raises if the restored core is not bit-exact
m.mount_arm("stages-1-4/s1_perspective")       # one member alone (the "alone" readings)

A group mounts its members in the order they were trained, always on, one after another at every block, with no mixer between them. default_arm in from_pretrained (or in config.json) mounts a group as the model loads. save_pretrained refuses while an arm is mounted, so an arm can never be written into the core's weight file.

Solo arms

Each stage also has an arm that was trained alone on the same frozen core, for when one stage's text is all that matters. A solo arm holds its whole stage by itself.

armits stage's textstage text, arm off → on (bpb)web text (bpb)stage items in the stage's own formstage items in a new formseeds
solo/s1_perspectiveone small event retold from each side: seen from a person, said to them, told about them0.554 → 0.047 (a second seed: 0.047)+0.000355% → 89%46% → 42%two
solo/s2_conceptkinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ0.721 → 0.075 (a second seed: 0.075)+0.000130% → 100%68% → 70%two
solo/s3_rulesif-then rules over made-up words, followed step by step to what follows0.930 → 0.180 (a second seed: 0.179)+0.000316% → 82%7% → 22%two
solo/s4_arithsmall arithmetic worked out in text1.234 → 0.267 (a second seed: 0.267)+0.000234% → 83%24% → 21%two
solo/s5_causalshort cause-and-effect text: what happened and why0.660 → 0.043 (a second seed: 0.043)+0.0006not measurednot measuredtwo
solo/s6_tryfailan attempt, its failure and the revised attempt0.673 → 0.057 (a second seed: 0.057)+0.0003not measurednot measuredtwo
solo/s7_mixedthe mixed stage's diet: the earlier stages' text beside ordinary prose0.783 → 0.715 (a second seed: 0.717)+0.0000not measurednot measuredtwo
solo/s8_registerthe same content in different registers of speech0.491 → 0.021 (a second seed: 0.021)+0.0003not measurednot measuredtwo
python
m.mount_arm("solo/s1_perspective")  # one solo arm; mounting another swaps it

Solo arms are for use one at a time. Each was trained to leave ordinary web text alone, but nothing taught it to stay out of the other stages' text. When the four solo arms of stages 1 to 4 were switched on together, three of the four stage texts read worse than with no arm at all. To have several stages on at once, mount the group: its arms were trained together for exactly that.

What an arm is

Each arm places one module after every decoder block. At a site it projects the block's output into 16 slots of dimension 8, reads them through a 16-atom aleph address (a closed-form signed read, never a selector), and adds a gated patch back into the residual stream:

slots = proj(x).view(B, n, 16, 8) mhat = sumk sinh(uk) Ak / sumk cosh(uk), u = (xhat . A) / tau x = x + sigmoid(gate) * consume(mhat)

13.7M parameters per arm over the 32 blocks. Every arm was trained with a quiet term: beside each chunk of its own stage's text it saw a chunk of text that is not its own and was penalised for changing the model's predictions there. For a group arm that text was ordinary web text and a partner arm's stage text, which is why the group can be left on and why no member harms a partner's stage. For a solo arm it was web text only.

Detach is exact

Mounting swaps the model's block list for wrapped blocks and keeps the originals. detach_arm() puts the originals back and re-runs a fixed probe: the restored logits must equal the pre-mount fingerprint bit for bit, or it raises. Before release, every packaged row was checked to produce logits identical to the adapter library the arms were trained with (amoe-lora).

The special-token control plane

Thirteen ids that valid UTF-8 can never produce carry structure and are trained:

idtokenmeaning
0xFFDOCdocument boundary (taught from step 0)
0xFE / 0xFDUSER / MODELturn openers (taught in the final chat phase)
0xFCENDuniversal block close
0xFBSYSsystem block opener
0xF7+bMODEregister tag (+1 ASCII byte)
0xF5+bESC254 extended slots
0xFA 0xF9 0xF8 0xF6 0xC0 0xC1THINK DATA SEP CUE RESreserved/instrument

Chat format: [SYS] text [END] [USER] text [END] [MODEL] text [END]. The frame is unforgeable (encoded text cannot contain a special). m.say() renders it for you; the chat frame was taught in the final 4.0B bytes, so treat it as a young capability.

Plain generation

python
prompt = "The history of astronomy begins"
ids = torch.tensor([list(prompt.encode("utf-8"))])
out = m.generate(ids, max_new_tokens=200, do_sample=True, top_p=0.95)
print(bytes(int(i) for i in out[0]).decode("utf-8", errors="replace"))

Weight files

fileprecisionwhat it is
model.safetensorsbf16the final weights (step 245,674); the default, and the core every arm was trained on
model.fp32.safetensorsfp32the trainer's master weights at the same step; load with variant="fp32"
arms/<group>/<arm>.safetensorsfp32one file per arm of a group, as trained
arms/solo/<arm>.safetensorsfp32one file per solo arm, as trained
arms/index.jsonthe arm table as data

Training ran under bf16 autocast over fp32 master weights. The bf16 file is those masters rounded to bf16, tensor for tensor. Each file's header names its precision and step. The arms were fitted on the bf16 file, so that is the core their measurements belong to.

Training

64.4B bytes on two RTX 5090s (data-parallel, 262,144 bytes per step): wikitext warmup (0.3B) → fineweb-edu (20.9B) → a nine-stage early-life curriculum, s0–s8 (35.2B: narrative, perspective, concepts, rule-chains, arithmetic, causal, try-fail, mixed, register) → a two-phase anneal (4.0B of distribution shift without the chat frame, then 4.0B with it). Muon + pure Adam split, flat LR, bf16 autocast, zero loss spikes and the gradient clip never reached across the entire run. Every boundary shipped a report (held-out bpb, the toggle ledger, the chat frame, a probe suite, the rank profile): the reports, every checkpoint, the resume states, the TensorBoard runs and the pinned training code live in alephllm-mini-beatrix-training under mini-beatrix-3/.

Final validation: 0.9861 bpb on the fineweb-edu holdout. The web pretraining phase closed at 0.9838; the curriculum's stage diets moved the web holdout as high as 1.15, and the anneal brought it back.

The arms. During curriculum stages 1–4 an arm was trained beside the trunk for each stage; those arms stayed nearly empty, because the trunk took each stage's text in before its arm could, and they were detached at step 148,000 (their files are in the training repo under mini-beatrix-3/arms/, bound to those earlier trunk states). The arms in this repository were fitted afterwards on the finished, frozen core: as always-on groups and as solo arms. The four-arm group in two steps (first the arms of stages 1 and 2, in a run that added the four stages' text one stage at a time with every attached arm training under one pure Adam, 800 steps per stage; then, with those two held fixed and on, a new arm for stage 3 and after it a new arm for stage 4, 800 steps each). The eight-arm group by continuing that run: with the four held fixed and on, a new arm for stage 5, then 6, then 7, then 8, each attached over every arm before it and trained for 4,000 steps in their presence. Every group arm was kept quiet on ordinary web text and on the stage text of every other arm in its group. Each solo arm was trained alone for 4,000 steps (two seeds), kept quiet on web text; and, as an experimental third group, with the eight held fixed and on, a ninth arm for image captions trained over them for 4,000 steps on the caption pack, kept quiet on web text and on the eight's stage text.

Lineage

Code: AbstractEyes/alephllm (this repo vendors the model files; alephllm 0.10.5) and AbstractEyes/amoe-lora (the arm library; arms.py here is a self-contained re-expression of its runtime). Siblings: mini-beatrix-2s (237M, 20 blocks, the previous rung), mini-beatrix-2.5s (the 2s core with its arms) and mini-beatrix-1 (112M, the 3-hub hybrid).

Licence

The weights are non-commercial: CC BY-NC-SA 4.0, because the training text includes CHILDES child-directed speech (CC BY-NC-SA, via TalkBank) and Wikipedia / WikiText (CC BY-SA). The code files, the model design and the full training process are MIT (AbstractEyes/alephllm); relay.py and arms.py are adapted from amoe-lora under Apache-2.0 (NOTICE). The full terms and the list of training sources: LICENSE.md.