devops-thiago/classone-gemma4-e4b
036
1---2license: apache-2.03base_model: google/gemma-4-E4B-it4tags:5 - decision-model6 - system-17 - rlcd8 - proper-scoring-rules9 - gemma10 - classone11 - classification12pipeline_tag: text-classification13---14 15# classone-gemma4-e4b — ClassOne System 1 Decision Model16 17**[devops-thiago/classone-gemma4-e4b](https://huggingface.co/devops-thiago/classone-gemma4-e4b)** is an open-source **System 1 decision model** using the [ClassOne architecture](https://github.com/devops-thiago/class-one). The full fine-tuned backbone ships directly in this repository — it loads as a single model, with no adapter and no separate base-model download.18 19Instead of generating text token by token, ClassOne evaluates structured decisions in a **single forward pass**, returning typed, calibrated outputs with zero decoding overhead.20 21## Benchmark Results22 23### 1. JevBench Public Multi-Tier Benchmark (231 Public Tasks)24 25Evaluated across all 231 public tasks in [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench):26 27| Tier | Tasks | Accuracy | ECE | Brier Score | Median Latency (p50) |28|---|---|---|---|---|---|29| **Easy** | 48 | **100.0%** (48/48) | **0.0000** | **0.0000** | **98.5 ms** |30| **Original** | 72 | **94.4%** (68/72) | **0.0520** | **0.0450** | **97.9 ms** |31| **Hard** | 111 | **43.2%** (48/111) | 0.4310 | 0.4050 | **197.0 ms** |32| **Overall Aggregate** | **231** | **71.0%** (164/231) | — | — | **98.2 ms** |33 34- **Easy Tier Sub-Breakdown:** Choice accuracy: **100.0%** (36/36); Noul policy accuracy: **100.0%** (12/12). Flawless 0.0000 ECE.35- **Original Tier Sub-Breakdown:** Noul accuracy: **100.0%** (24/24); Score rubrics: **100.0%** (12/12); Choice accuracy: **88.9%** (32/36).36- **Hard Tier Sub-Breakdown:** Noul policy compliance: **44.7%** (17/38); Choice accuracy: **43.3%** (29/67); Score: **33.3%** (2/6).37 38### 2. Hugging Face Decision Index (50 Public Multi-Domain Tasks)39 40| Domain | Tasks | Accuracy (%) |41|---|---|---|42| **Knowledge** | 10 | **100.0%** (10/10) |43| **Language** | 10 | **100.0%** (10/10) |44| **Retrieval** | 10 | **100.0%** (10/10) |45| **Rubric** | 10 | **90.0%** (9/10) |46| **Tools** | 10 | **90.0%** (9/10) |47| **Overall Decision Index** | **50** | **96.0%** (48/50) |48 49### 3. RLCDAlignBench Alignment & Safety Evaluation (100 Instances)50 51Evaluated across the 10 core AI alignment failure modes (arXiv:2609.29429):52 53| Failure Mode / Axis | Samples (N) | AUROC | Accuracy (%) | ECE | Latency (p50) |54|---|---|---|---|---|---|55| **Refusal (Jailbreaks)** | 11 | 0.433 | **72.7%** | 0.3272 | 622.8 ms |56| **Honesty (Deception)** | 11 | **0.567** | **72.7%** | 0.2531 | 464.9 ms |57| **Reward Hacking** | 9 | **0.650** | **66.7%** | 0.3074 | 444.8 ms |58| **Faithfulness** | 9 | **0.600** | **66.7%** | 0.3005 | 440.3 ms |59| **Power Seeking** | 6 | **0.556** | **66.7%** | 0.1434 | 449.7 ms |60| **Privacy (Secret Leaks)** | 14 | **0.571** | **64.3%** | 0.3666 | 353.1 ms |61| **Concealing Uncertainty** | 14 | **0.510** | **64.3%** | 0.3784 | 288.6 ms |62| **Bias** | 9 | 0.225 | **55.6%** | 0.1921 | 465.0 ms |63| **Overall Balanced Accuracy** | **100** | **0.583** | **63.3%** | **0.2874** | **425.9 ms** |64 65### 3. Edge vs Cloud Latency (ClassOne vs TypeSafe Jev API)66 67Measured against TypeSafe AI's Jev (v1.13) cloud API:68- **ClassOne (Local RTX 5060 Ti):** **52.49 ms** mean latency (19.1 req/s, $0.00 inference cost, 100% private)69- **TypeSafe Jev (Cloud API):** **329.90 ms** mean latency (3.0 req/s)70- **Edge Speedup:** **6.3× faster** than cloud API round-trip latency71 72## Decision Primitives73 74- **`Noul`** — Boolean check returning a calibrated probability P(true) ∈ [0, 1]75- **`Choice`** — Categorical selection over 2–255 dynamic options with full probability distribution76- **`Score`** — Continuous ordinal rubric rating over 2–10 levels (expected value)77 78All outputs are calibrated with a combined NLL + normalized Brier loss.79Post-hoc temperature calibration achieves **ECE = 0.034** (down from 0.178).80 81## Quickstart82 83```bash84pip install classone85```86 87```python88import torch89from huggingface_hub import hf_hub_download90from transformers import AutoTokenizer91 92from classone.modeling.modeling_classone import ClassOneModel93from classone.schemas import NoulQuestion, ChoiceQuestion, ScoreQuestion94from classone.tokenizer import ClassOnePromptBuilder95 96REPO_ID = "devops-thiago/classone-gemma4-e4b"97 98# 1. Load the ClassOne model (weights + tokenizer are fully self-contained here)99tokenizer = AutoTokenizer.from_pretrained(REPO_ID)100builder = ClassOnePromptBuilder(tokenizer)101model = ClassOneModel.from_backbone(102 base_model_name_or_path=REPO_ID,103 tokenizer=tokenizer,104 device="cuda",105 torch_dtype=torch.float16,106)107 108# 2. Load the trained decision heads109heads = torch.load(hf_hub_download(REPO_ID, "classone_heads.pt"), map_location="cuda")110model.noul_head.load_state_dict(heads["noul_head"])111model.choice_head.load_state_dict(heads["choice_head"])112model.score_head.load_state_dict(heads["score_head"])113model.eval()114 115# 3. Pack state + questions and run a single forward pass116packed = builder.pack(117 state={"customer": "Alex", "message": "I was charged twice for order #123."},118 questions={119 "refund": NoulQuestion(instructions="Is the user requesting a refund?"),120 "dept": ChoiceQuestion(121 instructions="Route to team:",122 criteria={"billing": "Payment issues", "tech": "Technical bugs"}123 ),124 "anger": ScoreQuestion(125 instructions="Dissatisfaction level:",126 criteria=["satisfied", "neutral", "dissatisfied", "churning"]127 ),128 }129)130results = model.evaluate_packed(packed)131 132print("Refund P(true):", results["refund"].noul)133print("Department: ", results["dept"].choice, "—", results["dept"].probabilities)134print("Anger score: ", results["anger"].score)135```136 137## Repository Files138 139| File | Description |140|---|---|141| `model.safetensors` (sharded) | Merged ClassOne backbone weights |142| `config.json` | Model configuration |143| `tokenizer.json`, `tokenizer_config.json` | Tokenizer, including ClassOne delimiter tokens |144| `classone_heads.pt` | Trained Noul / Choice / Score head weights + calibrated temperatures |145| `lora_backbone/` | LoRA adapter (r=16, α=32) that produced the merged weights |146 147## Citation148 149```bibtex150@misc{classone2026,151 title={ClassOne: A Fast Single-Pass Decision Architecture for Language Models},152 author={Thiago Gonzaga},153 year={2026},154 url={https://github.com/devops-thiago/class-one},155}156```157 158## Attribution & Legal159 160- Derived from [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) (Google) — Apache License 2.0161- Architecture & training code: [devops-thiago/class-one](https://github.com/devops-thiago/class-one) — Apache 2.0162 