Team Ai
Modelpublic

abhi6168/ekvachan-decoder-flash

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
1likes
Model Card

ekVachan-decoder-flash

LoRA adapters (r=16, alpha=32, 96 target modules) for Qwen/Qwen3.5-0.8B — the real-time arm of ekVachan, an open, self-hostable alternative to TypeSafe AI's Jev System-One decision model. Given a state and a set of typed questions (choice, score, noul), it returns calibrated probabilities in one non-autoregressive forward pass (a restricted-logit read over single-token letter codes), not a generated response.

Same mechanism and API as the full-size abhi6168/ekvachan-decoder (Qwen3.5-4B), trained through the same four-stage curriculum at roughly 5x the typical serving speed.

Full source, training pipeline, and evidence trail: github.com/asp616848/better-jev-for-all.

What's in this repo

Four LoRA adapters over the same base model, one per curriculum stage, hot-swapped at serving time by serve.inference.RoutingDecoderModel:

  • —ekvachan-decoder-qwen-benchcorpus/ — stage 1, text-only
  • —ekvachan-decoder-qwen-vision/ — stage 2, vision-capable; also answers text-only requests (the reference server's default text adapter)
  • —ekvachan-decoder-qwen-wide/ — stage 3, wide-option continue-training
  • —ekvachan-decoder-qwen-stage3/ — stage 4, atari-grok endgame (final)

Each subfolder carries its own manifest.json (training config, real measured accuracy/Brier/ECE) alongside the adapter weights — treat that file as the source of truth over this card if they ever disagree.

Real, verified numbers

Third-party benchmarks, zero training exposure, measured against the actual serving class (RoutingDecoderModel):

BenchmarkItemsAccuracy
JevBench (231/231 items)231/23156.71%
jabr-v2 (944/944 items)944/94467.37%
ViZDoom Defend the Center (von rubric)8 episodes3.0 kills avg
ViZDoom Health Gathering (von rubric)8 episodes34.3s survival avg

Typical latency (p50) on real JevBench-shaped requests through the production serving class with CUDA graphs on: ~24ms — versus ~112ms for the 4B model on the same workload.

Usage

Not a plain AutoModelForCausalLM.generate() model — it requires the restricted-logit read mechanism. Use serve.inference.RoutingDecoderModel from the main repo directly rather than reimplementing this:

python
from serve.inference import RoutingDecoderModel

# Point ./checkpoints at a local clone of this repo, then:
model = RoutingDecoderModel()
result = model.predict_choice(
    "The customer says the package never arrived.",
    ["refund", "replace", "escalate"],
)
print(result["choice"], result["probabilities"])

Or try it with zero setup in the ekVachan Space.

License

Apache-2.0, same as the base model and the main repo.