Team Ai
Modelpublic

Jwuthrich/selfjev-4b-vision-4bit

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
1likes32downloads
Model Card

SelfJev-4B Vision, 4-bit (NF4)

Structured decisions from text. A 4B backbone, quantized to run on an 8 GB GPU.

SelfJev-4B Vision with its LoRA merged into Qwen3.5-4B and the weights quantized to 4-bit (NF4) with bitsandbytes (NF4 4-bit with bf16 compute). It answers yes/no, one-choice, scored and multi-select questions over the options you supply and returns probabilities, through the SelfJev engine and API. The checkpoint is 3.3 GB (the full merged model is about 9.3 GB).

Quickstart

Needs an NVIDIA GPU (CUDA) and selfjev 0.4.0 or later.

bash
pip install "selfjev[serve,gpu,quant]"
selfjev serve --quantized-model Jwuthrich/selfjev-4b-vision-4bit

--quantized-model takes this repository (or a local folder) in place of --adapter. The first start also downloads the Qwen3.5-4B tokenizer; the vision tower is only fetched from the base model if you send an image. The HTTP API is the same as the full model's: API guide.

Quality

Scored on all 3,657 questions of SelfJev Decision Bench through the native tree engine on an A10G whose memory was capped at 7.2 GiB (the usable memory of an 8 GB card).

BuildDecision Bench (3,657)Text decisions (1,991)AI response review (946)Compact challenge (720)
bf16, full model95.9%96.1%92.5%100%
4-bit (NF4) (this repository)95.0%94.7%92.0%100%
8-bit (sibling build)95.3%95.2%92.0%100%

The suites are authored with AI-checked labels and have informed research decisions, so read them as a development benchmark. The full-model rows are the vision release's own reports; the compact bf16 figure was re-run on the same A10G. Reports: reports/quant.

Memory and speed

Peak GPU memory for one question and the median time of one request (A10G, --max-batch-tokens 4096, selfjev 0.4.2; memory is what the tensors allocate):

Text length4-bit (NF4)bf16 (full model)
2,048 tokens3.4 GiB, 454 ms8.3 GiB, 434 ms
8,192 tokens4.3 GiB, 1,565 ms9.2 GiB, 1,545 ms
16,384 tokens5.6 GiB, 3,193 ms10.4 GiB, 3,159 ms

This is the build for an 8 GB card: texts up to 16K tokens fit with room to spare, and speed matches the full model on this GPU.

Limits

  • —Scored on text only. Images use the base model's vision tower in bf16 (not quantized, loaded on the first image); the quantized path has not been scored on images.
  • —Texts above 16K tokens were not validated (the model was trained on texts up to 16K).
  • —Quantization costs accuracy: about 1 point on the pooled benchmark and 1.4 on text decisions, mostly in multi-select questions where the exact set of answers must match. The numbers above are the measurement; there is no calibration step to recover it.
  • —Same license and intended use as the full model; built from Qwen3.5-4B (Apache 2.0).

How this was made

selfjev quantize --adapter weights/selfjev_4b_vision --bits 4bit: the adapter is merged into the pinned Qwen3.5-4B (851bf6e) in bf16, then loaded with bitsandbytes NF4 4-bit with bf16 compute and saved. Source: Jwuthri/SelfJev.