mohit67890/imajev-4b
<p align="center"> <img src="assets/logo-light.png" class="block dark:hidden" width="300" alt="imajev"> <img src="assets/logo-dark.png" class="hidden dark:block" width="300" alt="imajev"> </p>
<h3 align="center">Decisions for real-world cases.</h3>
<p align="center">Small open models that read the photos, records and text a business already has and answer in the options you set, with a probability on each and an explicit <i>can't tell</i>. Your system acts when it is sure and hands the rest to a person.</p>
<p align="center"> <a href="https://huggingface.co/spaces/mohit67890/imajev"><img alt="Try the live demo" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/open-in-hf-spaces-md.svg"></a> <a href="https://github.com/mohit67890/imajev"><img alt="GitHub" src="https://img.shields.io/badge/GitHub-mohit67890%2Fimajev-111111?style=flat-square&logo=github"></a> <a href="https://mohit67890.github.io/imajev/"><img alt="Website" src="https://img.shields.io/badge/website-checked%20examples-555555?style=flat-square"></a> <a href="https://huggingface.co/datasets/mohit67890/imajev-bench"><img alt="ImajevBench" src="https://img.shields.io/badge/benchmark-ImajevBench-555555?style=flat-square"></a> <img alt="Apache-2.0" src="https://img.shields.io/badge/licence-Apache--2.0-111111?style=flat-square"> </p>
<p align="center"><b>imajev-4b</b> is the recommended default of the family · other sizes: <a href="https://huggingface.co/mohit67890/imajev-2b">imajev-2b</a> · <a href="https://huggingface.co/mohit67890/imajev-9b">imajev-9b</a><br> <a href="https://huggingface.co/spaces/mohit67890/imajev"><b>Live demo</b></a> · <a href="https://mohit67890.github.io/imajev/">Website</a> · <a href="https://github.com/mohit67890/imajev">Code and results</a> · <a href="https://mohit67890.github.io/imajev/report/">Technical report</a></p>
<p align="center"><a href="https://benchmarkheaven.com/jev-models"><img src="assets/ranks-light.png" class="block dark:hidden" width="820" alt="#1 of 91 on JevBench v1.4.2.2 (scored 27 Sep 2026); #1 of 49 on Image JevBench v0.1.3 (released 28 Sep 2026); #3 of 56 on DecisionBench (eng, v1), 28 Sep 2026."><img src="assets/ranks-dark.png" class="hidden dark:block" width="820" alt="#1 of 91 on JevBench v1.4.2.2; #3 of 56 on DecisionBench"></a></p>
<p align="center"><b>Independent results</b> (screenshots of the official leaderboards, 28 Sep 2026)</p> <p align="center"><a href="https://benchmarkheaven.com/jev-models"><img src="assets/proof-jevbench-score.png" width="760" alt="JevBench v1.4.2.2 Composite Score: 1 Imajev-4B 67.4, 2 Plumb-4B 65.8, 3 decider-4b v2 64.1, 4 Jev 1.13.0 63.3"></a></p> <p align="center"><a href="https://benchmarkheaven.com/image-jev-bench"><img src="assets/proof-image-jevbench.png" width="760" alt="Image JevBench v0.1.3 composite: 1 Imajev-4B 76.39, 2 Jev-Omni 73.10, 3 NeoHorse Jev 4B 71.94"></a></p> <p align="center"><a href="https://huggingface.co/spaces/Hanno-Labs/decision-bench-leaderboard"><img src="assets/proof-decisionbench.png" width="760" alt="DecisionBench (eng, v1): 1 bosun-v3.1-1.7b 87.29, 2 bosun-v3.1-0.6b 83.20, 3 imajev-4b 79.65, ahead of glm-5.3-flash, jev-1.13, deepseek-v4.1-flash, gpt-5.6-luna"></a></p> <p align="center"><sub>JevBench v1.4.2.2, scored 27 Sep 2026 (<a href="https://benchmarkheaven.com/jev-models">board</a>). Image JevBench v0.1.3, released 28 Sep 2026: #1 of 49, 76.39, ahead of Jev-Omni (<a href="https://benchmarkheaven.com/image-jev-bench">board</a>). DecisionBench (eng, v1): #3 of 56 models, 79.65, ahead of GLM-5.3 Flash (320B), Jev 1.13, DeepSeek V4.1 Flash (552B) and GPT-5.6 Luna; the two above are the benchmark team's own models (<a href="https://huggingface.co/spaces/Hanno-Labs/decision-bench-leaderboard">leaderboard</a>).</sub></p>
<p align="center"><img src="assets/request-listing.png" class="block dark:hidden" alt="imajev-4b checks a listing against its photo: listing.color says red, the photo shows beige shoes; the model names listing.color at 0.999 and the app holds the listing"><img src="assets/request-listing-dark.png" class="hidden dark:block" alt="imajev-4b checks a listing against its photo: listing.color says red, the photo shows beige shoes; the model names listing.color at 0.999 and the app holds the listing"></p>
imajev-4b is the recommended default: 83.9% on ImajevBench against the 9B's 82.1% at under half the size, and the only size trained in the phase-3 hard-data stage.
What sets it apart
<p><img src="assets/highlights.png" class="block dark:hidden" alt="Five highlights: a photo read against your record; two photos, one decision; a trained can't tell; open, small and local; Jev's contract, now with images"><img src="assets/highlights-dark.png" class="hidden dark:block" alt="Five highlights: a photo read against your record; two photos, one decision; a trained can't tell; open, small and local; Jev's contract, now with images"></p>
- A photo read against your record. Checks a photo against your own fields and names the one that is wrong. Trained on 72k photo-vs-record and two-photo decisions.
- Two photos, one decision. A reference and a target in the same request: shipped against returned, a known-good part against the one on the line.
- *A trained can't tell.* Every answer carries a probability for
unknown, so the app can stop instead of guessing. - Open, small and local. Apache-2.0, MLX on a Mac or PyTorch on one GPU; photos and customer data never leave your network.
- Jev's contract, now with images. TypeSafe's Jev request and response (
POST /v1/systemone), plusimages,unknown_probabilityandabstained. Jev itself is text-only and hosted; its state limit (32k tokens) is larger than imajev's (32 KB).
One request, every answer typed
The exact script we ran against imajev-4b and its output (rounded, usage shortened); 1.15 s on a Mac Studio (four option orders averaged, calibration file applied). Swap the adapter for this size and the request is unchanged.
import json, requests
URL = "http://127.0.0.1:8765/v1/systemone"
listing = {
"title": "Men's suede boat shoes",
"color": "red",
"product_type": "shoe",
}
questions = {
"contradicted_field": {
"type": "choice",
"instructions":
"Which field of `listing` does this photo contradict?",
"criteria": {
"listing.color": None,
"listing.product_type": None,
"none of these": "the photo agrees with every field",
},
},
"color_matches": {
"type": "noul",
"instructions":
"The product in the photo matches `listing.color`.",
},
"type_matches": {
"type": "noul",
"instructions": "The photo shows the kind of product "
"given in `listing.product_type`.",
},
}
request = {"state": {"listing": listing}, "questions": questions}
with open("listing.jpg", "rb") as photo:
r = requests.post(URL, files={"image": photo},
data={"request": json.dumps(request)})
print(json.dumps(r.json(), indent=2))<details><summary>Result</summary>
{
"model": "imajev-4b",
"answers": {
"contradicted_field": {
"type": "choice",
"choice": "listing.color",
"probabilities": {
"listing.color": 0.95,
"listing.product_type": 0.006,
"none of these": 0.043
},
"confidence": 0.919,
"unknown_probability": 0.007,
"abstained": false
},
"color_matches": {
"type": "noul",
"noul": 0.082,
"unknown_probability": 0.022,
"abstained": false
},
"type_matches": {
"type": "noul",
"noul": 0.989,
"unknown_probability": 0.004,
"abstained": false
}
},
"usage": {
"total_ms": 1152.9,
"input_tokens": 224
}
}</details>
<p><img src="assets/request-ticket.png" class="block dark:hidden" alt="A support ticket answered in one text-only request: department, urgency and frustration"><img src="assets/request-ticket-dark.png" class="hidden dark:block" alt="A support ticket answered in one text-only request: department, urgency and frustration"></p>
Automate what is clear, route the rest
<p><img src="assets/automation.png" class="block dark:hidden" alt="At a 90% threshold imajev-4b decides 58% of ImajevBench questions automatically, 97.5% of them correctly"><img src="assets/automation-dark.png" class="hidden dark:block" alt="At a 90% threshold imajev-4b decides 58% of ImajevBench questions automatically, 97.5% of them correctly"></p>
imajev-4b on the 279 ImajevBench test questions (photos, records and text; 21 whose honest answer is can't tell), raw probabilities, scored with the benchmark's own rule:
The rest go to a person. The benchmark is built to be hard; measure on a few hundred of your own cases before choosing a threshold. Other sizes at 90%: 2B 38% automated at 95.3% right, 4B 58% at 97.5%, 9B 70% at 91.8%.
Checked demos. Every clickable combination in the five playground apps (business checks, text only, wardrobe, stylist, tracing pad) was run on imajev-4b (four option orders) and compared with the right answer: 130 of 145 pass without the calibration file, 118 with it. Only passing combinations are shown as demos; the misses are listed in reports/scenarios/.
How it was made
About a million training decisions across the family, in four stages, for about $676 of rented GPU time for the whole project. The 4B was trained on stages 1 and 2 in one run (867k decisions: the 504k human-labelled set, 296k labelled by our 9B and 66k photo-vs-record and two-photo decisions), then on about 23k hard questions kept only when open-weight teachers agreed, then a soft-target continuation on 39,515 rows carrying Qwen3.6-35B-A3B's full probability distributions (with the strict slice of the Eikos decisions set (caiovicentino1/eikos-decisions, CC-BY-4.0; attribution and per-source licences in docs/eikos-decisions-usage.md) and 5k replayed image decisions). The shipped adapter is the weight-space average of two adapters: the hard-question adapter and that continuation. Every teacher is open-weight; no Jev outputs, paid-API outputs or JevBench items were used.
This adapter
This repository holds the 4B adapter: the recommended default tier — the best accuracy per millisecond in the family. It is a LoRA (rank 64, alpha 128) on the language layers of Qwen3.5-4B (revision 851bf6e8) plus a 256-code decision readout (255 option codes and unknown), in PEFT format at the root and in MLX format under mlx/. It is the last checkpoint of the phase-3 run: the previous release (a rank-16 weight-space average) expanded to rank 64 and trained for two rounds on the decisions that release got wrong. Code, server and evaluation harness: https://github.com/mohit67890/imajev. Other tiers: https://huggingface.co/mohit67890/imajev-2b (latency), https://huggingface.co/mohit67890/imajev-9b (quality); both are still the previous-generation adapters.
Technical specification
Training path. One trainer for every stage (PyTorch + PEFT): cross-entropy on the readout logits (soft targets where a record carries a distribution), AdamW with weight decay 0, linear warm-up then cosine decay to 10% of the peak rate, gradient clipping 1.0, seed 0, 4 GPUs. The soft-target stage adds a rationale loss (weight 0.3, at most 192 tokens) and permutes the options of every question.
Data this size saw.
- First run: 866,854 decisions: 504,000 stage-1 decisions with their original labels (36 licence-admitted sources); 296,482 stage-2 decisions labelled by the 9B, with
unknowntargets capped at 15%; 66,372 photo-vs-record and two-photo decisions. 17.15% of its training targets areunknown; no base-model blend. - Stage 3, round 1: 14,112 training records: kept teacher questions (9,368 of 13,386 kept on two-answerer agreement) plus the training share of 8,532 human reasoning items from 10 licensed sets.
- Stage 3, round 2: 7,812 training records: 3,598 new (4,852 of 8,097 kept on three-answerer agreement) + 4,214 replayed from round 1.
- Stage 4: 39,515 records: the stage-3 teacher questions relabelled with Qwen3.6-35B-A3B's probability distributions (thinking mode), 9,880 new hard, judge and programmatic questions, the strict slice of the Eikos decisions set (10,570 rows, open-weight teachers only) and 5,000 replayed image decisions.
- Stage 5 (phase 3), round 1: 125,424 rows: 58,000 hard text and 32,427 hard image decisions the previous release got wrong, mined from a 210,565-candidate pool (our generators, public datasets, earlier pools, trap variants), labelled by Qwen3.6-35B-A3B (thinking) with distribution targets and kept only after an independent review (Kimi-K2.5 plus a 220-item blind human-style review; families over 5% estimated label error dropped), including 20,270 constructed chart, document, inventory, safety, geometry and screenshot decisions with answers by construction; 19,998 earlier text and 14,999 earlier photo decisions replayed. 22% of targets are
unknown. - Stage 5, round 2: 26,494 rows: the 13,247 labelled items the round-1 model still failed plus 13,247 replayed.
Compute. $499.07 of rented GPU time on RunPod through stage 3, about $177 for stage 4 (all three sizes) and, for stage 5 (4B only), $177 of training (8×H100, 6 h 20 min incl. the pilot and the full evaluation) plus about $414 of mining, teacher labelling and review (RunPod and Azure, open-weight teachers): about $1,270 for the whole project, every run included.
Full specification: https://github.com/mohit67890/imajev/blob/main/docs/technical-specification.md
Results (2026-09-26; every number reproducible from the repo's reports/)
Every number in this table was produced in one evaluation on 2026-09-26 (pod ctr4dy9bzom5ji, 8×H100) with the previous release re-measured under the identical protocol; the full table for all eight phase-3 checkpoints, with item counts and gate results, is reports/phase3/train-results/benchmarks.md in the repo. The paired cluster test of this version against the previous one on ImajevBench (89 evidence clusters) is pending; the previous release beat its untuned base by +11.8 points [+5.8, +18.0], p = 0.0006. JevBench is text-only; imajev's image capability shows only on ImajevBench and in use. The official JevBench board adds 308 sealed items and a four-axis score run by its maintainer: there this adapter (revision c9e5f132, served with --rotations 1 --calibration calibration.json) is #1 of 91 on JevBench v1.4.2.2 (scored 27 Sep 2026; JevBench Score 67.37: Intelligence 52.2, Calibration 80.4, Speed 90.6, Cost 59.7; Jev 1.13.0 63.29). Source: https://benchmarkheaven.com/jev-models.
Calibration
calibration.json (schema 1.0) applies one temperature (1.305) to every question type × option-count bucket, fitted by negative log-likelihood on 150 template-generated JevBench-style items (none from JevBench). calibration-rot4.json is the same fit for the 4-rotation serving mode. A per-type fit on the flagged half of our held-out set was tried and rejected: it lowers hard-tier ECE but raises the pooled ECE over all public JevBench items (0.059 single / 0.039 rot4 against a 0.03 guard). Temperature scaling never changes an answer, only its probability. unknown offsets are 0.
Checked through the evaluation server (ECE, uncalibrated → with calibration.json); the off-distribution rows were measured on the previous version with its own temperature and have not been re-run for this version:
On photo-only verification the previous version's raw probabilities were already calibrated and the temperature over-softened them; if your traffic is mostly photo-against-record checks, serve without --calibration or fit your own temperature on a held-out sample.
Serving
git clone https://github.com/mohit67890/imajev && cd imajev
python3.11 -m venv .venv && . .venv/bin/activate
pip install -e ".[serve,mlx]" # Apple silicon; elsewhere: pip install -e ".[serve,torch]"
python scripts/download_model.py --model 4b
hf download mohit67890/imajev-4b --local-dir adapters/imajev-4b
# Mac (MLX)
PYTHONPATH=src:scripts python scripts/playground/server.py --model-bundle artifacts/model-qwen4b.json \
--adapter adapters/imajev-4b/mlx --rotations 4 --calibration adapters/imajev-4b/calibration-rot4.json --model-name imajev-4b --port 8765
# Linux / CUDA (PyTorch + PEFT)
PYTHONPATH=src:scripts python scripts/playground/server.py --backend torch --model-bundle artifacts/model-qwen4b.json \
--adapter adapters/imajev-4b --rotations 4 --calibration adapters/imajev-4b/calibration-rot4.json --model-name imajev-4b --port 8765Then POST /v1/systemone with a Jev-shaped request. One forward pass per question: p50 96 ms raw on one H100 for a JevBench hard item, serially; 350 ms with the --rotations 4 (four option orders averaged, +0.9 hard) and calibration.json used for the numbers above (shared pod, under load).
Training data and provenance
Synthetic documents and typed questions written by Qwen3.6-27B, answered independently by Qwen3.6-27B (thinking), gpt-oss-20b and, in the last part of the hard-question stage, Qwen3.6-35B-A3B (thinking); a question is kept only when every answerer agrees with the intended answer. In the soft-target stage the same questions were relabelled with Qwen3.6-35B-A3B's probability distributions (thinking mode), joined by 9,880 new hard, judge and programmatic questions, the strict slice of the Eikos decisions set (10,570 rows, open-weight teachers only) and 5,000 replayed image decisions; the shipped adapter is the weight-space average of the hard-question adapter and that continuation's best checkpoint. Plus licensed image datasets and human-written states from earlier training stages (see the repo's datasheets). No JevBench items (8-gram lint), no outputs from Jev or any paid API. All teachers are open-weight, Apache-2.0.
Limits
Single-pass: no reasoning at inference, so multi-step arithmetic and answer-quality judging trail reasoning models (a frozen Qwen3.6-35B-A3B with thinking scores 97% on JevBench hard at seconds per decision); the phase-3 stage moved our own hard held-out sets by 19 to 28 points but left JevBench hard unchanged within noise. Over-confident without calibration.json. On the previous release's 14-item unknown-gold check this version abstains on 11 (the previous release on all 14) while its false-abstention rate on answerable items stays at 0.5%: it is slightly less conservative on borderline "cannot tell" cases, and that ship gate was overridden for this release. Two-image comparisons are the weakest visual task (41.8% on real pairs, measured on an earlier imajev-4b); shelf inventory counts are the weakest constructed image family (67.6%). English only. The 2B and 9B tiers are still the previous generation.
