JohnP1/d1a-e4b-mlx-q8
D1A-E4B v0.6 · MLX 8-bit
JohnP1/d1a-e4b v0.6 for Apple Silicon. The build is the same as v0.5's:
- linear layers and token embeddings 8-bit (group 64), per-layer embeddings 4-bit;
- pointer head fp32;
- calibration temperature 1.95 in
d1a_config.json; - about 6 GB on disk, plus 1 GB of photo, voice and video encoders.
Only the first weight shard and the pointer head differ from v0.5: the per-layer embeddings (the second shard) and the encoders are byte-identical. The scores below were measured on this build.
D1A is a small open decision model. One document and a set of typed questions go in (yes/no, choice, score), and a calibrated probability for every option comes out, in one forward pass, with no generated text. It speaks the TypeSafe System One API (POST /v1/systemone), so the TypeSafe SDK and existing clients work against it.
v0.6 labels pull requests the way they are labelled today. On 487 pull requests created after all of its training, severity goes from 74.7% to 79.1% (+4.3 points) and change type from 86.7% to 90.6% (+3.9) against v0.5, and P1 recall returns to v0.4's 0.35 (v0.5: 0.22). It also gains on long documents (+2.0), agent-factory routing (+3.4) and held-out transfer (+1.4), with lower calibration error on most suites. It costs hard decisions (−2.7) and change type on the older frozen PR test set (−4.1), so it ships as a release exception (below). v0.5 stays available. See Results.
Release exception. v0.6 does not pass D1A's release rule against v0.5. scripts/decide.py vetoes it (V1) on hard decisions: hard-v1 −2.7 points [−4.9, −0.5], 61 fixes against 90 breaks, past the −2.0 limit for a card suite. It also loses change type on the older frozen PR test set, −4.1 [−6.0, −2.3] (19 fixes, 58 breaks), and pooled over all 21 lines it is level with v0.5, −0.20 [−0.98, +0.52]. It ships anyway, by the owner's decision of 2026-10-10: the model's live job, the one model on the owner's Mac mini, is labelling current pull requests, and on PRs newer than all of its training it is clearly better. The likeliest reading of the disagreement is that the project's labelling conventions have changed over time and v0.6 follows the current ones. Stay on v0.5 (JohnP1/d1a-e4b-mlx-q8@v0.5) if you rely on hard-v1-style reasoning or on the older PR-label conventions.
What's inside
A LoRA adapter (rank 16 on the attention and MLP projections: q, k, v, o, gate, up, down) and a small pointer head (256-dimensional) on google/gemma-4-E4B at revision 411aa17b. Each version continues training the previous one, so v0.6 contains every stage below:
Before training, the v0.6 mix was screened against 47 evaluation partitions, the 487 newer PRs below included: none of their documents or PRs is in it.
Calibration: a single temperature, T = 1.95, fitted on pooled held-out rows (decision-v7 calibration split and PR-labeling development set; calibration error 0.093 before, 0.015 after, out of fold), stored in d1a_config.json. v0.6 carries no per-use-case temperatures yet, so a request with "use_case": "routing" is read at 1.95 too; a refitted routing temperature is planned as v0.6.1.
Results
Pull requests newer than all training
487 hermes-agent pull requests created after every PR in v0.6's training, held back before v0.6 was trained (the test was registered in advance). All three models are MLX 8-bit at their served temperatures, with the serving context. Differences are paired over PRs, with 95% bootstrap intervals:
Against v0.4, v0.6 is +1.6 [−0.8, +4.1] on severity and +4.1 [+1.4, +7.0] on change type. A cheaper alternative, v0.5 with outcome memory fitted on the same recent labels, reached 89.3% on change type but only 76.4% on severity (+1.6 [−1.0, +4.3]), with P1 recall 0.17, and failed its own bar.
Every suite, against v0.5
Held-out data only. Both builds are MLX 8-bit (this repository) at their served temperatures (v0.5 1.78, v0.6 1.95), scored with d1a.eval.benchmark through D1A's quality gate and scripts/decide.py. Differences are paired over questions, with 95% bootstrap intervals:
Pooled over all 21 lines, v0.6 − v0.5 = −0.20 [−0.98, +0.52]. Latency is unchanged: v0.6 over v0.5 on short requests 1.0009 [0.9991, 1.0035], within run-to-run noise (measured interleaved).
¹ These PR sets are older than the 487 above. v0.5's numbers here are its 8-bit build at T 1.78 over the full sets, so they differ a little from v0.5's own card, which scored PR labels on the bf16 PyTorch checkpoint. ² The PRs the owner's labeling job saw on 2026-10-06 and 07, labelled by the owner on the playground's review page.
Against Kev-4B's model card (fp32, development splits; an approximate comparison): hard decisions 78.6, developer tools 73.9, long documents 89.1, transfer 81.7. v0.6 is level with Kev-4B on long documents.
Calibration: v0.6's single temperature fits its calibration sets (0.015 out of fold), and its calibration error is lower than v0.5's on 15 of the 22 suites the gate scores (most on transfer-v4, agent factory and wanli-v2). It is higher on the other 7, most on dates (0.028 → 0.053), typesafe-v1 (0.104 → 0.129) and JGLUE (0.026 → 0.042).
Photos, voice and video
media/ holds Gemma 4 E4B's own vision and audio encoders (bf16, ~1 GB, unchanged from the base: the LoRA adapter only touches the text model; byte-identical to v0.5's). With them d1a.serving.serve answers questions about a photo, a voice clip or a video with this same model.
- The endpoint is
POST /v1/systemone/media: a /v1/systemone request plus{"media": {"type": "image" | "audio" | "video", "data": <base64>}}. - The encoders turn the media into soft tokens, which the 8-bit language model reads right after
<state>. - A video is read as 16 timestamped frames; its sound is not used.
- The encoders are fetched on the first such request; a plain download skips them.
No media in training (zero-shot). On the playground's photo, voice and video examples, v0.6 changes 1 of 14 requests from v0.5 (one video example).
The questions it was trained for
PR labeling works best with these questions, worded exactly so, over a document of the form title / author / stats / body / files (one - status path +added/-deleted line per file):
type(choice): "Primary change type from files and body, not the title prefix." Options: type/bug, type/docs, type/feature, type/perf, type/refactor, type/security, type/test.blast(choice): "How far a mistake in this PR spreads in production." Options: review:blast-contained (one module), -moderate (one subsystem), -broad (shared helper or config), -massive (auth, permissions, or all paths).sev(choice): "How serious the problem this PR addresses is — not the risk of merging the diff as-is." Options P0 to P4.
The PR-labeler recipe has the exact wording, the document builder and the training and scoring scripts. The mix builder builds and screens mixes like v0.6's. Other questions about other documents work as in v0.2.
Run it
On Apple Silicon (the PyTorch checkpoint, for NVIDIA GPUs and CPUs, is JohnP1/d1a-e4b v0.6):
pip install "d1a[serve] @ git+https://github.com/jonpol01/d1a"
python -m d1a.serving.serve --run JohnP1/d1a-e4b-mlx-q8@v0.6 --port 8009or in-process: from d1a import D1A; D1A.load("JohnP1/d1a-e4b-mlx-q8@v0.6").decide(state, questions). For hard-v1-style reasoning or the older PR-label conventions, load @v0.5.
Limits
- Hard decisions are 2.7 points below v0.5 (68.8% against 71.5%; the release exception above), and 10 below Kev-4B (78.6).
- Change type on older PRs is worse: 85.3% against 89.4% on the frozen English test set, and 72.5% against 76.2% on the Japanese one (all questions, 92 PRs). It calls more PRs type/bug than v0.5 did. On current PRs it is better (above).
- Blast radius on the owner's own repositories is worse: 67% against 74% on 39 hand-checked PRs (1 fix, 4 breaks). v0.6 was not trained for blast radius; on the larger frozen test set it is 62% against 58%.
- Severity: on the 487 newer PRs, P2 recall falls (0.55 against v0.5's 0.79) while P3 and P0 rise; P0 rests on 7 PRs. Severity on the older development set is 1.8 points lower.
- One temperature for every question, and no use-case temperatures (routing's is planned for v0.6.1). Calibration error is higher than v0.5's on JGLUE, dates and typesafe-v1.
- Smaller losses against v0.5: decision-v7 −1.4, wanli-v1 −2.3, one of 45 hand-labelled routing questions.
- The recent PR labels come from one large open-source project and a few days of it; if its conventions move again, so will these numbers.
- Questions are in English; the documents can be English or Japanese.
Versions
Older tag names (v0.1-2epoch, v0.1.1-2epoch-calibrated, v0.2-hybrid) still work.
License and data
Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). Code: github.com/jonpol01/d1a, built on Kev by Jared Palmer (Apache-2.0). Japanese decision data derived from JGLUE by Yahoo Japan Corporation and Waseda University (CC BY-SA 4.0). Pull-request data from NousResearch/hermes-agent (MIT); the labeled PR dataset itself, including the recent PRs, is private.
