Team Ai
Modelpublic

prithivMLmods/JEV-9B-GGUF

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
3likes1.8kdownloads
Model Card

JEV-9B-GGUF

autotrust/JEV-9B is AutoTrust AI's first integrated System 1 + System 2 open model, built on a frozen, bit-identical Qwen3.5-9B backbone using a "Blocks of Experts" recipe: System 2 is ordinary text generation through the untouched base lm_head (70.7% HumanEval pass@1, identical to base Qwen3.5-9B with all 164 completions byte-identical), while System 1 is a small, detachable 40.2M-parameter LoRA plus a 24-slot fp32 decision head that answers typed noul (yes/no), choice (2–16 options), or score (0–5 scale) questions in a single forward pass, distilled from the closed, hosted TypeSafe Jev 1.13's own output distributions via the Apache-2.0 SargeDev/jev-distill-corpus-v3 corpus. On 25,376 Jev-labelled held-out rows, JEV-9B reaches a mean KL divergence of just ≈0.019 nats from the teacher's distributions (essentially indistinguishable at that resolution, including reproducing several of the teacher's known mistakes), 90.2% choice top-1 agreement, 0.994 noul AUROC, and an ECE of 0.0007 with no post-hoc correction needed, while also generalizing to unseen task families (KL 0.234, top-1 91.8% on out-of-distribution Open-Jev rows) and reaching 90–97% of the teacher's accuracy on an independent third-party benchmark with human gold labels. It is dramatically faster than the hosted API — a single decision takes ~90ms median versus 238–301ms for the hosted service, and one B200 GPU sustains ~15x the throughput — and both systems are served from one set of weights via a single vLLM engine, with a request routed to either path per-call; its larger sibling, autotrust/JEV-27B, trades some of this speed for closer teacher fidelity, better OOD transfer, and stronger System 2 generation (78.0% HumanEval). The model, its LoRA adapter, decision head, and training/evaluation reports are all released under Apache-2.0, and AutoTrust AI states it is an independent, unaffiliated reproduction sharing no code or weights with TypeSafe AI.

Model Files

File NameQuant TypeFile SizeFile LinkDescription
JEV-9B.BF16.ggufBF1617.9 GBLinkFull BF16 weights. Highest quality, largest file size.
JEV-9B.Q3KL.ggufQ3KL4.93 GBLinkLower quality but usable, good for low RAM availability.
JEV-9B.Q3KM.ggufQ3KM4.62 GBLinkLow quality.
JEV-9B.Q4KM.ggufQ4KM5.63 GBLinkGood quality, default size for most use cases, recommended.
JEV-9B.Q4KS.ggufQ4KS5.35 GBLinkSlightly lower quality with more space savings, recommended.
JEV-9B.Q5KM.ggufQ5KM6.47 GBLinkHigh quality, recommended.
JEV-9B.Q5KS.ggufQ5KS6.31 GBLinkHigh quality, recommended.
JEV-9B.Q6_K.ggufQ6_K7.36 GBLinkVery high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp