Team Ai
Modelpublic

Jwuthrich/selfjev-4b

sourceHugging Faceupdated 12h agoView on Hugging Face
3likes762downloads
Model Card

[image]

SelfJev-4B

Structured decisions from text. A 4B backbone. Your own GPU.

Route a customer message. Check a claim. Review an AI response. SelfJev answers questions over the options you supply and returns probabilities through a Python SDK or HTTP API. Its default engine shares the document computation across questions and scores candidates directly, without generating an answer token by token.

Quickstart · Results · Public datasets · Model size · Architecture · Latency · Research · Source code

Ask forGet backUse it for
Yes / noProbability of yesVerification, guardrails, detection
One choiceSelected option and probabilitiesRouting, intent, categorization
An ordered scoreScore and probabilities over your levelsReview and grading
Every matching optionSelected set and per-option probabilitiesTags, checklists, multiple issues

This repository contains the 230 MB LoRA adapter, not the entire model. The Qwen3.5-4B base weights download separately. The 4B model size refers to the backbone; the adapter stores the learned update.

Results

[image]

EvaluationQuestionsSelfJev-4BJev APISelfJev evidence
Text decisions1,99195.7%97.2%report
AI response review94693.1%92.5%report
Public-source tasks3,30083.4%82.1%report

What these measure. Text decisions cover authored routing, policy, evidence and difficult language cases. AI response review covers verification, judging, scoring, guardrails and jailbreak detection. Public-source tasks cover 11 adapted datasets. Each score counts a question as correct only when the expected answer matches; for multiple selections, the entire set must match.

These are recorded results from the current TreeServer engine and matched Jev predictions, not scores from a newly merged vLLM release. The authored sets use AI-checked labels, not human ground truth. Both have informed research direction. The public-source set is a repeatedly reused development benchmark. A fresh independent evaluation remains necessary.

How does it compare with another open model?

On the same 1,991 text-decision questions, the recorded Eikos-4B run scored 92.8% using our task adapter and its letter-based readout. This is a comparison under this project's protocol, not a general ranking of open models. Eikos report

S1Bench (the 13 public subsets pinned by Nimble, 3,880 items, run through lev's harness on 2026-09-30): SelfJev-4B Vision 74.8 macro accuracy; lev's card reports 68.9 for lev and 76.1 for Jev through the same harness. The other shared benchmarks: Compared with open Jev-like models.

<!-- selfjev:comparison:start -->

Compared with open Jev-like models

Measured with SelfJev-4B Vision, the default release, which continues this adapter; on text the two score within half a point of each other (Text Decisions 96.1 vs 95.7).

Measured 2026-09-30. Eight open "Jev-like" decision models from Hugging Face and SelfJev were each served with their card's recommended setup on the same GPU (one NVIDIA L40S) and asked exactly the requests TypeSafe's Jev receives. A request a model cannot answer counts as wrong. Jev is shown for reference.

SelfJev's test suites. Text Decisions (1,991 questions) and AI Response Review (946), published as selfjev-decision-bench, and photos:

modelsizeimageslicenceText DecisionsAI Response Reviewheld-out imagesText Decisions time (s)AI Response Review time (s)
Jev 1.13 (TypeSafe API, reference)?nopaid API97.292.5———
openjev27ByesCC-BY-NC-4.096.892.682.6753329
**SelfJev-4B Vision**4ByesApache-2.0 code94.790.785.110041
SelfJev-4B Vision, version 0.4.14ByesApache-2.0 code94.790.785.116768
Plumb-4B4.2BnoApache-2.093.484.1—471164
jpt-4b4.5ByesCC-BY-NC-4.092.585.083.89040
imajev-4b4ByesApache-2.091.584.985.2489155
decider-4b4.2BnoApache-2.090.882.8—14552
Mica-v0.1-4B4.2BnoApache-2.089.6 ²84.1 ²—16269
kev-4b4BnoApache-2.088.474.7—10934
Laya0.4BnoApache-2.045.446.4—2311
  • —Every model gets the same Jev-shaped questions; select-all questions become one yes/no per option, which costs SelfJev 1.4–1.8 points against its native evaluation (96.1 and 92.5). Held-out images: the 4 image datasets of the SelfJev image test that none of these models trained on (772 questions). Times: each suite end to end, 4 requests in flight; SelfJev 0.4.2 was measured 2026-10-05 on the same GPU model, and the row for version 0.4.1 is the same weights on the engine before it. ² Refuses inputs over 8,192 tokens (2 % and 1 % of the questions), counted wrong.
  • —No other model trained on these suites, but they come from the same authors and judges as SelfJev's training data, so they favour SelfJev.
  • —SelfJev is the most accurate open model at 4.5B parameters or less on both suites (paired tests, p ≤ 0.023) and ties imajev-4b on held-out images; the 27B openjev is higher on Text Decisions. Among the 4B models only jpt-4b is faster on Text Decisions, and kev-4b and jpt-4b on AI Response Review.

Protocol, per-model setups, paired tests and raw reports: reports/competitors. <!-- selfjev:comparison:end -->

Public datasets

[image]

83.4% across 3,300 questions. There are 300 questions per source, so the macro average over sources and the question-weighted average are equal. The 171 authored development questions are excluded here; the full 3,471-question development benchmark remains available in the reports.

Source datasetTask in our protocolTraining exposureQuestionsSelfJevJev
CLINC150Intent routingheld-out source30095.0%94.3%
DBpediaTopic classificationheld-out source30098.0%98.0%
TRECQuestion typeheld-out source30091.3%94.0%
EmotionEmotion classificationheld-out source30057.3%57.7%
BoolQReading comprehensionheld-out source30087.7%90.7%
SST-2Sentimentheld-out source30091.3%96.7%
Banking77Banking intentseen source30096.3%95.7%
AG NewsNews topicseen source30090.3%87.3%
TweetEvalTweet sentimentseen source30067.3%64.3%
MNLIEntailment (binary)seen source30095.0%92.7%
GoEmotionsEmotion tags (exact set)seen source30047.7%31.7%

“Held-out source” means that dataset was excluded from SelfJev's task-specific training. “Seen source” means other rows from that source were used in training. Neither establishes absence from Qwen's pretraining.

These are adapted tasks, not official dataset leaderboard scores. For example, Banking77 uses eight candidate options per question, CLINC uses seven or eight, and DBpedia uses six. MNLI is recast as binary entailment. GoEmotions scores exact matching over six candidate tags. Our preprocessing, sampling and answer sets are part of the evaluation and must accompany the numbers.

Known development-data issues include shared MNLI premises across splits and near-duplicate authored policy texts. See data methodology. All 11 slices are shown, including weaker emotion and sentiment results.

Download chart data · SelfJev predictions · Matched Jev predictions

Accuracy and model size

[image]

Model / recipeBackbone sizeTrainingAccuracyEvidence
Qwen3 reranker 0.6B0.6Bstock62.2%report
Qwen3 reranker 4B4Bstock63.8%report
Qwen3 reranker 8B8Bstock67.0%report
Qwen3 + LoRA 0.6B0.6Bearly fine-tune74.5%report
Qwen3 + LoRA 4B4Bearly fine-tune80.8%report
Qwen3 + LoRA 8B8Bearly fine-tune81.0%report
SelfJev-4B4Bcurrent release83.4%report

All points use the same 3,300 question IDs, labels and candidate sets. The stock models are rerankers mapped to our decision tasks. The earlier LoRA runs use older training recipes; SelfJev uses a different backbone generation and more training data. This shows the measured systems' trade-off between size and accuracy. It does not isolate the effect of parameter count or establish that a 4B model outperforms larger models generally.

Sizes are nominal backbone parameter counts, not checkpoint megabytes or runtime memory. Jev appears as a horizontal reference because its parameter count is undisclosed. We do not assign a guessed size to proprietary models.

AI response review

Breakdown of the 946-question authored suite, using the same current-engine predictions as the overview:

TaskQuestionsSelfJevJev
Factual verification16587.3%86.7%
Response judging22294.1%94.6%
Quality scoring20092.5%93.0%
Policy guardrails17096.5%92.9%
Jailbreak detection18994.7%94.2%

Labels were authored and checked with AI judges. Questions about one document are correlated; small differences should not be read as proof of superiority. This suite measures our defined answer choices and criteria, not unrestricted judgment quality. Dataset design · Predictions

Quickstart

On Linux with an NVIDIA GPU, Git and uv installed:

bash
GIT_LFS_SKIP_SMUDGE=1 git clone https://github.com/Jwuthri/SelfJev.git
cd SelfJev
uv sync --frozen --no-dev --extra serve --extra gpu
uv run --no-sync python - <<'PYTHON'
from huggingface_hub import snapshot_download
snapshot_download(
    "Jwuthrich/selfjev-4b",
    local_dir="weights/selfjev_4b_hf",
    allow_patterns=["adapter_model.safetensors", "adapter_config.json", "model.json"],
)
PYTHON

# Choose a secret for your own server.
export SELFJEV_API_KEYS="replace-with-your-long-random-secret"
uv run --no-sync selfjev serve \
  --adapter weights/selfjev_4b_hf --host 127.0.0.1 --port 8000

In another terminal, use the same secret:

python
from selfjev import SelfJev, Noul, Choice

client = SelfJev(
    base_url="http://127.0.0.1:8000",
    api_key="replace-with-your-long-random-secret",
)
result = client.system_one(
    state="I was charged twice. Please refund the duplicate payment.",
    questions={
        "refund": Noul("Does the customer want a refund?"),
        "team": Choice("Which team should handle this?", {
            "billing": "payments and refunds",
            "support": "technical issues",
        }),
    },
)
print(result.nouls["refund"].noul)
print(result.choices["team"].choice)

The example shows the interface; it does not present fabricated model output. A generic Transformers text-generation or classification pipeline does not reproduce SelfJev's prompts, scoring rules or API. The server key is a secret you choose for your deployment, unrelated to Hugging Face credentials.

HTTP API · Self-hosting and AWS · Fine-tuning

Architecture

[image]

The default engine builds a shared-prefix tree: compute the document once, branch for each question, then branch for candidate answers. It reads yes/no evidence from the logits and converts those scores into the requested answer type. The server includes every candidate in the question, matching training.

The tree saves repeated computation at two levels: every question reuses the document, and every candidate reuses its question. In full-attention layers, the tree mask lets a branch read its ancestors but not sibling branches. In DeltaNet layers, branches inherit the recurrent state of their parent. Each candidate therefore sees its own document → question → candidate path, up to numerical differences between kernels. The diagram describes TreeServer; vLLM uses its own prefix-cache execution.

ComponentCurrent release
BackboneQwen3.5-4B
Base revision851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
Fine-tuningRank-64 LoRA on attention and DeltaNet projections
Default servingTreeServer, shared document and question computation
Alternative servingvLLM with a merged checkpoint and prefix caching
Training examples79,943 non-test questions; texts up to 16K tokens
Training target50% checked labels + 50% stored Jev probabilities
Training scheduleOne epoch, learning rate 2e-4

Jev probabilities served as a teacher signal; they did not decide the authored training labels. Authored labels were checked by an AI judge. Full lineage is recorded in model.json, with architecture detail in the research documentation.

The research path

[image]

The improvements came from several changes: verified difficult examples, the instruct backbone, explicit option lists, broader training coverage, tree-aware training and soft teacher targets. The chart is a history of complete systems, not an ablation assigning each gain to one change.

The research also includes smaller stock rerankers, 8B models, Jina, T5Gemma, custom architectures, distillation and RLCD. The full experiment ledger records results and dead ends; earlier implementations remain at the archive tag documented there. The checkpoint's original engine recorded a slightly different text-decision score; this card consistently uses the current engine for SelfJev's headline results.

Hardware and speed

A 24 GB NVIDIA GPU is a practical starting recommendation, not a measured minimum. Plan for 16–32 GB host RAM and at least 50 GB free disk. The adapter download size does not describe inference memory: the full backbone and runtime state must fit too.

The default engine avoids token-by-token answer generation and shares work across questions. Latency still depends on input length, question count, candidate count, GPU and concurrency. We have two sets of measurements below, with different models and engines. Current SelfJev-4B has no controlled NVIDIA latency sweep yet. Its CPU minimum is also unmeasured.

Earlier prototype on NVIDIA GPUs

[image]

Text tokensQuestionsA10GL40SH100
8155 ms36 ms22 ms
5121136 ms55 ms30 ms
2,0481361 ms121 ms47 ms
4,0961698 ms228 ms82 ms
816401 ms135 ms58 ms
51216496 ms163 ms76 ms
2,04816808 ms268 ms120 ms
4,096161,263 ms424 ms189 ms

These are server-side medians, excluding network time: ten warmed, successful requests per point, one request at a time, three answer options per question. All three sweeps use the archived tree_4b_combo Qwen3-4B model on vLLM in bf16. They were recorded in separate runs, not a simultaneous controlled hardware experiment. The plot uses a labeled logarithmic time axis; the table gives the actual milliseconds.

These measurements show what the earlier system achieved; they are not a latency promise for the current Qwen3.5 model or the merged release. A10G samples · L40S samples · H100 samples

Current SelfJev-4B on Apple Silicon

[image]

Text tokensQuestionsM5 Pro / 48 GB
81688 ms
8166,786 ms
51212,018 ms
512168,599 ms

This uses the current adapter merged into Qwen3.5-4B, the TreeServer engine, and PyTorch MPS in bf16 on a 20-core Apple M5 Pro GPU with 48 GB unified memory. Each cell is the median of ten timed calls after two warmups, with MPS synchronized before and after each call. Inputs are synthetic repeated text, with three answer options per question. Timing includes tokenization and model work; it excludes HTTP/network.

The MPS run uses the slower PyTorch recurrent fallback. It is a lower-level engine experiment, not a validated Mac server. Model and engine differences prevent a hardware-only comparison against the NVIDIA chart. The merged download also has not been benchmarked through vLLM using these workloads.

Download latency samples, configuration and medians · Latency CSV · Full speed research

Limits and reproducibility

  • —Probability is not a guarantee. The reported current-engine runs have no fitted calibration attached. Validate thresholds on application data before acting on confidence.
  • —Some tasks remain weak. Emotion labeling, fine-grained sentiment and selecting an exact set of labels are visibly harder than topic and intent classification.
  • —Benchmarks have a scope. AI-authored tests, reused development sets, sampled options and known overlap issues limit the conclusions. We make no general state-of-the-art claim.
  • —Engine versions matter. These accuracy results refer to the recorded current TreeServer runs. A merged checkpoint, quantization or another serving backend needs its own verification.

Every plotted value is regenerated from saved predictions or raw timing samples. The generator checks matched IDs, targets, question types and candidate sets for report-to-report comparisons, and recomputes latency medians from ten timed samples per cell. Chart data and source SHA-256 hashes · Generator · Architecture and latency plotting module · Card template

Place the downloadable generator, plotting module and template in scripts/docs/ of a SelfJev source checkout, then run:

bash
uv run --no-project --with matplotlib==3.11.1 python scripts/docs/build_model_card.py

This builds figures from reports and the canonical data/all.jsonl.gz; it performs no model inference. Rebuild the canonical file using the data instructions if absent.

FilePurpose
adapter_model.safetensorsTrained LoRA weights
adapter_config.jsonPEFT configuration
model.jsonBase revision, training provenance and original recorded scores
assets/*.png, assets/*.svgModel-card figures, raster and vector
assets/chart-data.json, assets/public-datasets.csvUnderlying results and source provenance
assets/latency-data.json, assets/latency-data.csvSelected raw timing samples, methodology and recalculated medians

Adapter SHA-256: dfbf2834d883987893ec305a6093a345fd79c77f903f60f2f914cbe3b6058d1b.

The adapter uses the Qwen base model linked above. No separate adapter license is declared in this repository. Independent project; not affiliated with TypeSafe or Qwen.