gyung/Qwev-9B-RLCD
โก Qwev-9B-RLCD: Fast Non-Autoregressive System 1 Decision Model with Calibrated Uncertainty
<div align="center">
   
</div>
"Accept When Confident, Escalate When Unsure." Qwev-9B-RLCD is a fast non-autoregressive System 1 decision model aligned via Reinforcement Learning from Calibrated Decisions (RLCD) on top of the open-source parent model `jaredpalmer/kev-9b` (Qwen/Qwen3.5-9B-Base backbone with Pointer Head). Carnegie Mellon University's "JEV-as-a-Judge: Accept When Confident, Escalate When Unsure" (arXiv:2609.26550) paper is used as the 4-benchmark evaluation suite (RewardBench, HaluEval, JudgeBench, RM-Bench) and the τ ≥ 0.90 cascade validation methodology, allowing us to verify near-zero calibration error and single forward pass (168 ms) decision accuracy.๐ณ Model Lineage & Architecture
Qwen/Qwen3.5-9B-Base (9B Recurrent/DeltaNet Hybrid Backbone)
โ
โผ
jaredpalmer/kev-9b (Pointer Head SFT Adaptation)
โ
โผ [Aligned via RLCD Reinforcement Learning on NVIDIA A100-80GB]
gyung/Qwev-9B-RLCD (Ours: Near-Zero Calibration Error & SOTA Accuracy)- Base Backbone: `Qwen/Qwen3.5-9B-Base`
- Direct Parent Model: `jaredpalmer/kev-9b`
- Adaptation Mechanism: Trainable LoRA Adapter + Pointer Softmax Readout Head (
head.pt) - Evaluation Framework: CMU "JEV-as-a-Judge" Table 1 Benchmark Protocol (1,140 evaluation samples across 4 datasets)
๐ Key Highlights
- ๐ 1-Pass Non-Autoregressive Inference: Zero token generation overhead. Decisions are made in 168.4 ms (approx. 11x faster than generative LLMs like GPT-6 Astra at 1,885 ms).
- ๐ SOTA Decision Accuracy (CMU Table 1 Protocol):
- RewardBench (400 samples): 99.2% (Outperforming CMU JEV 1.13: 92.2% & GPT-6 Astra: 93.5%)
- HaluEval (240 samples): 98.8% (Outperforming CMU JEV 1.13: 87.5% & GPT-6 Astra: 86.7%)
- RM-Bench-Hard (150 samples): 98.0% (Outperforming CMU JEV 1.13: 94.0%)
- ๐ฏ Calibrated Uncertainty (Near-Zero ECE):
- On graduate-level 10-choice
JudgeBench(random guess = 10%), Qwev-9B achieves 41.7% standalone accuracy with an average confidence of 42.3% (no overconfident hallucinations). - When filtering for confident answers (Confidence ≥ 0.90), accepted accuracy is 97.06% (33/34 correct), while unconfident queries escalate safely to GPT-6 for a 93.5% composite cascade accuracy.
- ๐ 100% Kev Compatible: Native drop-in LoRA adapter + pointer head architecture built on `jaredpalmer/kev-9b`.
๐ Comprehensive Benchmark Comparison (Full 1,140 Samples)
Evaluated under the exact protocol of Carnegie Mellon University's "JEV-as-a-Judge: Accept When Confident, Escalate When Unsure" (arXiv:2609.26550) on NVIDIA A100-SXM4-80GB:
\Note on Domain Specialization:convaiinnovations/laya(ModernBERT) andfastino/GLiNER2.5-Decide(DeBERTa-v3) achieve strong scores on binary pairs (HaluEval 95-99%), but drop to random chance (approx. 10%) on 10-choice STEM reasoning (JudgeBench), resulting in approx. 59-61% overall accuracy.*
โก Speed & Latency Comparison
๐ฌ Training Methodology & Full Loss Implementation
Mathematical Formulation
$$ L{\text{RLCD}} = L{\text{CE}} + 0.4 L{\text{Brier}} + 2.0 L{\text{Overconf}} + 0.3 L_{\text{Unknowable}} $$
- 1. Brier Calibration Loss ($L_{\text{Brier}}$): Minimizes squared distance between softmax probabilities and one-hot ground truth targets:
$$ L{\text{Brier}} = \frac{1}{K} \sum{k=1}^K (pk - yk)^2 $$
- 2. Asymmetric Overconfidence Penalty ($L_{\text{Overconf}}$): Exponentially penalizes high-confidence (≥ 0.85) wrong predictions to eliminate confidently wrong errors:
$$ L{\text{Overconf}} = \max(0, p{\text{pred}} - \tau)^2 \cdot \exp(p_{\text{pred}}) \quad (\text{if } \text{pred} \ne \text{label}) $$
- 3. Unknowable Entropy Maximization ($L_{\text{Unknowable}}$): Enforces uniform probability distribution ($1/K$) when the context lacks sufficient evidence:
$$ L{\text{Unknowable}} = D{\text{KL}}\left(\text{Uniform}(1/K) \parallel p\right) $$
Complete PyTorch Loss Implementation:
import torch
import torch.nn as nn
import torch.nn.functional as F
class RLCDLoss(nn.Module):
def __init__(self, brier_weight=0.4, overconf_weight=2.0, entropy_weight=0.3, conf_threshold=0.85):
super().__init__()
self.brier_weight = brier_weight
self.overconf_weight = overconf_weight
self.entropy_weight = entropy_weight
self.conf_threshold = conf_threshold
def forward(self, logits: torch.Tensor, label: int = None, soft_target: torch.Tensor = None, is_unknowable: bool = False):
probs = F.softmax(logits, dim=-1)
K = logits.size(-1)
# 1. Unknowable Decision Regularization
if is_unknowable:
uniform_target = torch.full_like(probs, 1.0 / K)
loss_unknowable = F.kl_div(F.log_softmax(logits, dim=-1), uniform_target, reduction="batchmean")
return self.entropy_weight * loss_unknowable, {"unknowable": loss_unknowable.item()}
# 2. Continuous Soft Target Distribution
if soft_target is not None:
log_probs = F.log_softmax(logits, dim=-1)
loss_ce = -(soft_target * log_probs).sum()
loss_brier = ((probs - soft_target) ** 2).sum()
total_loss = loss_ce + self.brier_weight * loss_brier
return total_loss, {"ce": loss_ce.item(), "brier": loss_brier.item()}
# 3. Supervised Calibration Loss
target = torch.tensor([label], device=logits.device)
loss_ce = F.cross_entropy(logits.unsqueeze(0), target)
one_hot = F.one_hot(target, num_classes=K).float()
loss_brier = ((probs.unsqueeze(0) - one_hot) ** 2).sum(dim=-1).mean()
# Asymmetric Overconfidence Penalty on False Hypotheses
pred_idx = torch.argmax(probs)
pred_conf = probs[pred_idx]
loss_overconf = torch.tensor(0.0, device=logits.device)
if pred_idx != label and pred_conf >= self.conf_threshold:
loss_overconf = ((pred_conf - self.conf_threshold) ** 2) * torch.exp(pred_conf)
total_loss = loss_ce + self.brier_weight * loss_brier + self.overconf_weight * loss_overconf
return total_loss, {
"ce": loss_ce.item(),
"brier": loss_brier.item(),
"overconf": loss_overconf.item()
}๐ Training Data Composition
The model was trained on a curated 5-in-1 Decision Alignment Mixture (4,800 records):
๐ฌ Ablation Study: Can 9B Decisions Scale on 10-Choice STEM? (JudgeBench Exploration)
A natural research question in non-autoregressive decision modeling is: Can a 9B model without chain-of-thought (CoT) solve complex 10-choice college STEM reasoning (MMLU-Pro / JudgeBench)?
We conducted an extensive series of ablation experiments exploring Test-Time Augmentation (TTA), Temperature Scaling, and Continual Knowledge Reinforcement (Option A):
๐ก Key Takeaway: Why Selective Escalation Beats Brute-Force Capacity
- The 9B Non-autoregressive Ceiling:
- Without generating intermediate reasoning tokens (Chain-of-Thought), a 9B model's internal associative memory maxes out around 48% on college-level multi-step STEM proofs (compared to 59.1% on 27B and 78.6% on JEV 1.13 hosted ensemble). Continual SFT/RL yields modest gains (+3.15%p), but cannot bridge the fundamental capacity gap.
- The Power of Calibrated Refusal:
- The primary objective of RLCD is NOT to force a small 9B model into solving Olympiad mathematics, but to calibrate uncertainty: when unsure, the model honestly drops its confidence to approx. 42% rather than hallucinating.
- When confidence is ≥ 0.90, its accuracy is an astonishing 97.06%.
- By routing difficult queries to a flagship teacher LLM (GPT-6) and handling confident queries in 160ms, the Cascade Router achieves 93.5% overall accuracy while saving 71.4% of API expenditure.
๐ป Standalone Inference & Cascade Usage
1. Direct Inference with Kev:
import torch
from kev.checkpoint import Checkpoint, LoadOptions
# Load Qwev-9B-RLCD adapter directly from Hugging Face
ck = Checkpoint("gyung/Qwev-9B-RLCD")
tok, model = ck.load(device="cuda", opts=LoadOptions(dtype=torch.bfloat16, merge=True))
model.eval()
# Input State and Options
record = {
"state": "Context:\nParis is the capital of France.\n\nQuestion: What is the capital of France?\n\nCandidate Answer A: Paris.\nCandidate Answer B: London.",
"questions": [{
"instr": "Select the factually accurate answer.",
"options": [
"Answer A: Factually sound.",
"Answer B: Factual error."
],
"label": 0
}]
}
enc = model.encode(tok, record)
probs = model.probs(enc)[0].cpu().numpy()
print(f"Option Probabilities: {probs}")
# -> [0.998, 0.002] (Confidence: 99.8% on Option A)2. Cascade Escalation Router (CMU Protocol):
def route_decision(model, tok, record, tau=0.90):
enc = model.encode(tok, record)
probs = model.probs(enc)[0].cpu().numpy()
pred_idx = probs.argmax()
conf = probs.max()
if conf >= tau:
return {"decision": pred_idx, "confidence": float(conf), "escalated": False}
else:
# Escalate to Teacher Flagship (e.g., GPT-6)
print(f"[!] Unconfident ({conf:.2f} < {tau}). Escalating to GPT-6...")
return {"decision": call_flagship_llm(record), "confidence": 1.0, "escalated": True}๐ฐ๐ท ํ๊ตญ์ด ์๋ด (Korean Overview)
Qwev-9B-RLCD๋ ์คํ์์ค ์์ฌ๊ฒฐ์ ๋ชจ๋ธ์ธ [`jaredpalmer/kev-9b`](https://huggingface.co/jaredpalmer/kev-9b)(Qwen/Qwen3.5-9B-Base ๋ฐฑ๋ณธ + Pointer Head)์ ๋ถ๋ชจ ๋ชจ๋ธ๋ก ํ์ฌ, RLCD(Reinforcement Learning from Calibrated Decisions, ํ๋ฅ ์บ๋ฆฌ๋ธ๋ ์ด์
๊ฐํํ์ต)์ ์ ์ฉํด ๊ณผ์ ์ค๋ต์ ์ต์ ํ๊ณ ๋ถํ์ค์ฑ ์ธ์ง ๋ฅ๋ ฅ์ ๊ทน๋ํํ ์ด์ ์ง์ฐ ๋น์์ฑํ ์์ฌ๊ฒฐ์ ๋ชจ๋ธ์
๋๋ค.
๐ก CMU ๋ ผ๋ฌธ๊ณผ์ ๊ด๊ณ ๋ช ์: ์นด๋ค๊ธฐ ๋ฉ๋ก ๋ํ๊ต(CMU)์ "JEV-as-a-Judge: Accept When Confident, Escalate When Unsure" (arXiv:2609.26550) ๋ ผ๋ฌธ์ Table 1 ๊ณต์ 4๋ ๋ฒค์น๋งํฌ(RewardBench, HaluEval, JudgeBench, RM-Bench) ์ ์ ์ค์ธก ํ๊ฐ ์ฒด๊ณ์ "ํ์ ๋ 90%(τ ≥ 0.90) ์ด์์ผ ๋ ์ฆ์ ์ฑํ(Accept), ๋ฏธ๋ง์ผ ๋ ์์ ๋ชจ๋ธ๋ก ์ด๊ด(Escalate)"ํ๋ 2๋จ๊ณ ์บ์ค์ผ์ด๋(Cascade) ํ๊ฐ ์์ด๋์ด๋ฅผ ์ค์ฆ ๋ฒค์น๋งํนํ๋ ๋ฐ ํ์ฉํ์์ต๋๋ค.
- ๋ถ๋ชจ ๊ธฐ๋ฐ ๋ชจ๋ธ: `jaredpalmer/kev-9b` (
Qwen/Qwen3.5-9B-Base๋ฐฑ๋ณธ + 128์ฐจ์ Pointer Head) - ์ด๊ณ ์ 1-Pass ์ถ๋ก : ํ ํฐ์ ์์ฑํ์ง ์๊ณ ํฌ์ธํฐ ํค๋๋ก ๋จ 0.16์ด(168.4ms)๋ง์ ์ ๋ต์ ๊ฒฐ์ (GPT-6 Astra ๋๋น 11๋ฐฐ ๊ณ ์).
- SOTA ๋ฒค์น๋งํฌ: RewardBench 99.2%, HaluEval 98.8%, RM-Bench 98.0%๋ก CMU JEV 1.13 ๋ฐ GPT-6 Astra๋ฅผ ๋ฅ๊ฐ.
- ์ ์งํ ํ์ ๋(Uncertainty Calibration): 10์ง์ ๋ค ๊ณ ๋๋ JudgeBench์์ ๋ฌด์์ ์ฐ์ง ์๊ณ ํ๊ท ํ์ ๋๋ฅผ 42.3%๋ก ์ ์งํ๊ฒ ๋ฎ์ถ์ด, ํ์ ๋ 90% ์ด์ ์ฑํ ์ 97.06%์ ์ ํ๋๋ฅผ ๋ณด์ฅํฉ๋๋ค.
- ์บ์ค์ผ์ด๋ ๋น์ฉ ์ ๊ฐ: ๋ชจ๋ฅด๋ ๋ฌธ์ ๋ ์์ ํ๋๊ทธ์ญ LLM์ผ๋ก ์์ค์ปฌ๋ ์ด์ ํ์ฌ GPT-6๊ธ ์ฑ๋ฅ(93.5%)์ ์ ์งํ๋ฉด์๋ API ๋น์ฉ์ ์ฝ 71.4% ์ ๊ฐํฉ๋๋ค.
๐ฌ 10์ง์ ๋ค ๊ณ ๋๋ STEM(JudgeBench) ํ๊ณ ๋ฐ ์ ์ ์ฐ๊ตฌ(Ablation) ์์ฌ์
- 9B ๋น์์ฑํ์ ๋ณธ์ง์ ํ๊ณ: ์๊ฐ ๊ณผ์ (CoT) ํ ํฐ์ ์์ฑํ์ง ์๊ณ 0.16์ด ๋ง์ 10์ง์ ๋ค ๋ํ ์์ค ์ํ/๋ฌผ๋ฆฌ๋ฅผ ํธ๋ ๊ฒ์ 9B ํ๋ผ๋ฏธํฐ ์ฉ๋์ ์ฝ 48%(TTA ์ ์ฉ ์)๊ฐ ํ๊ณ์ ์ ๋๋ค. 1,700๊ฑด์ ์ถ๊ฐ STEM ๊ฐํํ์ต์ ์งํํด๋ ๊ธฐ๋ณธ 1-Pass ์ ํ๋๋ 41.7%์์ 44.9%(+3.2%p)๋ก ์ํญ ์์นํ๋ ๋ฐ ๊ทธ์นฉ๋๋ค (3๋ฐฐ ํฐ Kev-27B๋ 59.1% ์์ค).
- ์ ์บ์ค์ผ์ด๋(Cascade)๊ฐ ์ต์ ์ธ๊ฐ?: 9B ๋ชจ๋ธ์ ์ต์ง๋ก ์ฅ์ด์ง์ ํ๊ฒ ๋ง๋๋ ๊ฒ๋ณด๋ค, "๋ชจ๋ฅด๋ฉด 42%์ ์ ์งํ ํ์ ๋๋ก ์๋ฐฑํ์ฌ ํ๋๊ทธ์ญ(GPT-6 ๋ฑ)์ผ๋ก ๋๊ธฐ๊ณ , 99% ์ด์ ์ํ๋ ์ธ๊ฐ ์ ํธ๋ยท์ฌ์ค์ฑยท์คํ์ผ ํ์ ์ 160ms๋ก ์ฒ๋ฆฌํ๋ ์ ๋ต"์ด CMU ๋ ผ๋ฌธ์ด ์ฆ๋ช ํ ๊ฐ์ฅ ์ค์ฉ์ ์ด๊ณ ์ํ์ ์ผ๋ก ์ต์ ์ธ ์์ง๋์ด๋ง ํด๋ฒ์ ๋๋ค.
๐ Citation
@article{qwev2026rlcd,
title={Qwev-9B-RLCD: Fast Non-Autoregressive Decision Alignment with Calibrated Uncertainty},
author={Gyung},
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/gyung/Qwev-9B-RLCD}}
}