Lucebox/DeepSeek-V4.1-Flash-ROCMFP23-GGUF
DeepSeek V4.1 Flash for the Lucebox engine (ROCMFP block formats)
Lucebox builds of DeepSeek V4.1 Flash for the Lucebox DS4.1 engine (AMD R9700 + Strix Halo). Routed experts in the engine's ROCMFP block formats; attention, shared experts and head in Q8_0; both Engram tables embedded. Calibrated with the published 262,144-token imatrix (smalinin/DeepSeek-V4.1-Flash-GGUF, converted in imatrix/).
Formats: Q2_0_ROCMFP2 (ggml type 107, 10 bytes per 32 weights, 2.5 bits, codebook {-1, 0, 1, 2} x a ue4m3 scale per 16 weights), Q3_0_ROCMFPX (type 104, 14 bytes per 32, 3.5 bits). fp2s files use type 107 with bit 7 of a half-block's scale byte meaning "mirror the codebook"; a reader that ignores that bit produces garbage, so use an engine build that honors it. KL is measured on 8,184 held-out fineweb-edu tokens against the unquantized model through the same reference forward pass (eval/ holds the teacher log-probs and the tokens). Recipe, converter patches and harness: Lucebox repo, lucebox_training/quantization.
IQ2 / IQ3 + MXFP8 builds
ds41-lucebox-final.gguf (recommended) and ds41-lucebox-minus10gib.gguf (10 GiB smaller): routed experts in IQ2_XXS / IQ3_XXS, dense projections in MXFP8 (type 117, needs an engine with MXFP8, see Luce-Org/lucebox#800). Correct answers on our internal hard bench; exact = plain greedy, no drafter.
