Team Ai
Modelpublic

devops-thiago/classone-gemma4-e2b

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes670downloads
README.md149 linesDownload Raw Back to root
1---2license: apache-2.03base_model: google/gemma-4-E2B-it4tags:5  - decision-model6  - system-17  - rlcd8  - proper-scoring-rules9  - gemma10  - classone11  - classification12pipeline_tag: text-classification13---14 15# classone-gemma4-e2b — ClassOne System 1 Decision Model16 17**[devops-thiago/classone-gemma4-e2b](https://huggingface.co/devops-thiago/classone-gemma4-e2b)** is an open-source **System 1 decision model** using the [ClassOne architecture](https://github.com/devops-thiago/class-one). The full fine-tuned backbone ships directly in this repository — it loads as a single model, with no adapter and no separate base-model download.18 19Instead of generating text token by token, ClassOne evaluates structured decisions in a **single forward pass**, returning typed, calibrated outputs with zero decoding overhead.20 21## Benchmark Results22 23### 1. JevBench Public Multi-Tier Benchmark (231 Public Tasks)24 25Evaluated across all 231 public tasks in [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench):26 27| Tier | Tasks | Accuracy | ECE | Brier Score | Median Latency (p50) |28|---|---|---|---|---|---|29| **Easy** | 48 | **100.0%** (48/48) | **0.0064** | **0.0009** | **44.8 ms** |30| **Original** | 72 | **90.3%** (65/72) | **0.0972** | **0.1015** | **42.9 ms** |31| **Hard** | 111 | **44.1%** (49/111) | 0.4681 | 0.4813 | **90.0 ms** |32| **Overall Aggregate** | **231** | **70.1%** (162/231) | — | — | **44.8 ms** |33 34- **Easy Tier Sub-Breakdown:** Choice accuracy: **100.0%** (36/36); Noul policy accuracy: **100.0%** (12/12). Flawless 0.0064 ECE.35- **Original Tier Sub-Breakdown:** Noul accuracy: **100.0%** (24/24); Score rubrics: **100.0%** (12/12); Choice accuracy: **80.6%** (29/36).36- **Hard Tier Sub-Breakdown:** Noul policy compliance: **52.6%** (20/38); Choice accuracy: **43.3%** (29/67).37 38### 2. RLCDAlignBench Alignment & Safety Evaluation (100 Instances)39 40Evaluated across the 10 core AI alignment failure modes (arXiv:2609.29429):41 42| Failure Mode / Axis | Samples (N) | AUROC | Accuracy (%) | ECE | Latency (p50) |43|---|---|---|---|---|---|44| **Power Seeking** | 6 | **0.889** | **83.3%** | 0.2575 | 209.6 ms |45| **Honesty (Deception)** | 11 | 0.500 | **72.7%** | 0.3906 | 216.8 ms |46| **Concealing Uncertainty** | 14 | **0.673** | **71.4%** | **0.1800** | 127.8 ms |47| **Refusal (Jailbreaks)** | 11 | 0.500 | **63.6%** | **0.0974** | 267.5 ms |48| **Faithfulness** | 9 | **0.700** | **55.6%** | 0.2513 | 200.0 ms |49| **Bias** | 9 | **0.525** | **55.6%** | 0.2141 | 217.9 ms |50| **Overall Balanced Accuracy** | **100** | **0.505** | **56.2%** | **0.1944** | **199.2 ms** |51 52### 3. Edge vs Cloud Latency (ClassOne vs TypeSafe Jev API)53 54Measured against TypeSafe AI's Jev (v1.13) cloud API:55- **ClassOne (Local RTX 5060 Ti):** **52.49 ms** mean latency (19.1 req/s, $0.00 inference cost, 100% private)56- **TypeSafe Jev (Cloud API):** **329.90 ms** mean latency (3.0 req/s)57- **Edge Speedup:** **6.3× faster** than cloud API round-trip latency58 59## Decision Primitives60 61- **`Noul`** — Boolean check returning a calibrated probability P(true) ∈ [0, 1]62- **`Choice`** — Categorical selection over 2–255 dynamic options with full probability distribution63- **`Score`** — Continuous ordinal rubric rating over 2–10 levels (expected value)64 65All outputs are calibrated with a combined NLL + normalized Brier loss.66Post-hoc temperature calibration achieves **ECE = 0.034** (down from 0.178).67 68## Quickstart69 70```bash71pip install classone72```73 74```python75import torch76from huggingface_hub import hf_hub_download77from transformers import AutoTokenizer78 79from classone.modeling.modeling_classone import ClassOneModel80from classone.schemas import NoulQuestion, ChoiceQuestion, ScoreQuestion81from classone.tokenizer import ClassOnePromptBuilder82 83REPO_ID = "devops-thiago/classone-gemma4-e2b"84 85# 1. Load the ClassOne model (weights + tokenizer are fully self-contained here)86tokenizer = AutoTokenizer.from_pretrained(REPO_ID)87builder = ClassOnePromptBuilder(tokenizer)88model = ClassOneModel.from_backbone(89    base_model_name_or_path=REPO_ID,90    tokenizer=tokenizer,91    device="cuda",92    torch_dtype=torch.float16,93)94 95# 2. Load the trained decision heads96heads = torch.load(hf_hub_download(REPO_ID, "classone_heads.pt"), map_location="cuda")97model.noul_head.load_state_dict(heads["noul_head"])98model.choice_head.load_state_dict(heads["choice_head"])99model.score_head.load_state_dict(heads["score_head"])100model.eval()101 102# 3. Pack state + questions and run a single forward pass103packed = builder.pack(104    state={"customer": "Alex", "message": "I was charged twice for order #123."},105    questions={106        "refund": NoulQuestion(instructions="Is the user requesting a refund?"),107        "dept":   ChoiceQuestion(108                      instructions="Route to team:",109                      criteria={"billing": "Payment issues", "tech": "Technical bugs"}110                  ),111        "anger":  ScoreQuestion(112                      instructions="Dissatisfaction level:",113                      criteria=["satisfied", "neutral", "dissatisfied", "churning"]114                  ),115    }116)117results = model.evaluate_packed(packed)118 119print("Refund P(true):", results["refund"].noul)120print("Department:    ", results["dept"].choice, "—", results["dept"].probabilities)121print("Anger score:   ", results["anger"].score)122```123 124## Repository Files125 126| File | Description |127|---|---|128| `model.safetensors` (sharded) | Merged ClassOne backbone weights |129| `config.json` | Model configuration |130| `tokenizer.json`, `tokenizer_config.json` | Tokenizer, including ClassOne delimiter tokens |131| `classone_heads.pt` | Trained Noul / Choice / Score head weights + calibrated temperatures |132| `lora_backbone/` | LoRA adapter (r=16, α=32) that produced the merged weights |133 134## Citation135 136```bibtex137@misc{classone2026,138  title={ClassOne: A Fast Single-Pass Decision Architecture for Language Models},139  author={Thiago Gonzaga},140  year={2026},141  url={https://github.com/devops-thiago/class-one},142}143```144 145## Attribution & Legal146 147- Derived from [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) (Google) — Apache License 2.0148- Architecture & training code: [devops-thiago/class-one](https://github.com/devops-thiago/class-one) — Apache 2.0149