Team Ai
Modelpublic

mindchain/imajev-9b-GGUF

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
3likes671downloads
Model Card

![license](https://huggingface.co/mindchain/imajev-9b-GGUF) ![vision](https://huggingface.co/mindchain/imajev-9b-GGUF) 8 GB limit-orange) quantizations temperature verified

What the image path actually does

<img src="https://huggingface.co/mindchain/imajev-2b-GGUF/resolve/main/assets/imajev-vision-scan.gif" width="900" alt="five images through the decision readout">

The same page, then the same page blackened. Same model, same question, same temperature. The distribution moves from A 0.482 to D 0.887 — the image reaches the forward pass.

This is the one test that separates a vision decision model from a text decision model that has been given a picture it cannot see: send an obviously different image and check whether the answer changes. Twenty seconds of work, and on this hardware it is what proved the Pascal vision tower runs (--no-mmproj-offload) and what caught every wrong turn on the way.

What the numbers below do and do not show

The distributions on this card were measured with the task in the state. That matters more than it sounds. Asked "does this page fulfil its task?" without the task, the same model gave nearly flat answers — spread 0.089 across twenty pages instead of 0.176 — because it was guessing, and guessing looks plausible. On this hardware a decision model can only judge a page against a task it has actually been given.

So: the image path works, and the numbers move when the question is answerable. What they do not do is agree with the rule grader. Measured on the same pages, the correlation is −0.38, which reads like a defect and is not one — it is the direction of the labels. Read option D rather than A and the ordering follows the rule:

rule coverage >= 0.7   D 0.230   "no" rare
rule coverage <  0.7   D 0.429   "no" common

No two graders agree, and now one of them at least looks.

Where these rank, and what that means here

Two leaderboards, two different questions. Both are public and both are frozen releases, so the numbers below can be checked.

Images — [Image JevBench v0.1.5](https://benchmarkheaven.com/image-jev-bench), 50 ranked systems:

#1  Wity-1 (API, base model undisclosed)   80.2
#2  Imajev-4B                              76.4   ← best open model
#6  JPT-4B                                 69.5   ← also in this portfolio
#7  imajev 2B                              68.7

Text — [JevBench v1.5.4](https://benchmarkheaven.com/jev-models/v1.5.4), 106 ranked systems:

#1   Jev 1.13.0 (API, closed)              80.0
#10  Imajev-4B                             70.8
     Plumb-4B                              ~65.8 (v1.4.2.1)

So the honest claim, and it is narrower than the one we started with: Imajev-4B is the strongest open image decision model measured, and it runs on an 8 GB card. Not the strongest overall — that is Wity-1, a hosted model we cannot inspect — and not the strongest text model, which is the closed Jev. For text at that level the open field is thin, which is why this collection is weighted toward the image side.

A caveat we measured ourselves: the composite weighs intelligence, calibration, speed and cost. On this hardware the published cost and speed figures do not apply — these are laptop numbers, and the Q4KM builds were never benchmarked on this task. The ranking says the decision quality is there; it does not say the quantizations preserve it.

Which size, and why

Your cardTakeWhy
4 GB[2B](https://huggingface.co/mindchain/imajev-2b-GGUF) at Q3KS, 0.95 GBthe only one that leaves room for context on a 4 GB card
8 GB[4B](https://huggingface.co/mindchain/imajev-4b-GGUF) up to Q80 · **[9B](https://huggingface.co/mindchain/imajev-9b-GGUF)** up to Q5K_M4B is rank #1 on JevBench; the 9B needs 7.44 GiB there including projector and KV cache
12 GB+[9B](https://huggingface.co/mindchain/imajev-9b-GGUF) up to Q8_0, 8.87 GBthe largest, and the one whose image path is still untested on our hardware

All three share the same decision method and the same readout. They differ in base size, in their calibration temperature, and in how much VRAM they need.

Whole collection: Imajev GGUF

If you want to build your own

One command, and five traps that cost a night:

bash
bash build_imajev_gguf.sh <merged-dir> <name> <out-dir>

The traps are the reason this repository exists. Qwen3.5 vision models do not convert with the obvious incantation:

  1. 1.Look the architecture up in the table, conversion/__init__.py, not in a filename. There is no qwen35.py.
  2. 2.`--mmproj` writes the projector and silently drops the weights — 0.62 GB instead of 3.52 GB.
  3. 3.The config announces an MTP layer the checkpoint does not contain. --no-mtp fixes the text path and is not defined for the vision one.
  4. 4.Patching `block_count` by hand creates a second bug. The original value was right; the model was incomplete.
  5. 5.The mmproj step fails silently unless preprocessor_config.json is copied in from the base model.

Plus one that is not in the script: the calibration temperature is not in the GGUF. Divide the option logits by it yourself, client-side.

How the three graders disagree

Measured on 20 pages with three independent graders:

rule (source vs task)   <-> Imajev-2B     rho = +0,086
rule (source vs task)   <-> JPT-4B         rho = +0,027
text decider (Tev1)     <-> image decider  rho = +0,299
two image deciders      <-> each other     rho = +0,265

No two graders agree. That is why this is a family and not one "best" model, and why the training pipeline composes them multiplicatively (R = R_test x S_sol x S_beh) instead of adding them: additive rewards make GRPO maximise the strongest component and sacrifice the weakest, which is exactly how the coverage score and the rendering score drifted apart here.

A fourth check belongs with them. Send a black image and a white one: if the distribution does not move, the model is reading text, not pixels. Twenty seconds of work, and it is what proves the Pascal vision tower runs at all.

Imajev-9B GGUF — the largest Imajev, and where the 8 GB line ends

GGUF weights that actually load. As of October 2026 there was no GGUF for Imajev anywhere on the Hub: mohit67890/imajev-2b, -4b and -9b ship safetensors only, the GGUF-tag search returns nothing, and the one third-party variant (zenmagnets/Imajev-4B-FP8-SM120) is FP8 for Hopper (SM120) and will not run on a consumer card.

This repository provides the two GGUF files plus the conversion recipe that was missing, for the model people with small VRAM actually want to run.

Part of the collection, together with the other two sizes: https://huggingface.co/collections/mindchain/imajev-gguf-vision-decision-models-that-actually-load

Files

Full standard quantization range, as bartowski and unsloth publish them.

FileSizefits 8 GB?
imajev-9b-Q3_K_S.gguf3.97 GB5.39 GiB — yes
imajev-9b-Q3_K_M.gguf4.31 GB5.73 GiB — yes
imajev-9b-Q3_K_L.gguf4.59 GB6.01 GiB — yes
imajev-9b-Q4_K_S.gguf4.98 GB6.40 GiB — yes
imajev-9b-Q4_0.gguf4.95 GB6.37 GiB — yes
imajev-9b-Q4_K_M.gguf5.24 GB6.66 GiB — yes
imajev-9b-Q5_0.gguf5.87 GB7.29 GiB — yes
imajev-9b-Q5_K_M.gguf6.02 GB7.44 GiB — yes
imajev-9b-Q5_1.gguf6.33 GB7.75 GiB — yes, barely
imajev-9b-Q6_K.gguf6.85 GB8.27 GiB — no
imajev-9b-Q8_0.gguf8.87 GB10.29 GiB — no
imajev-9b-mmproj-f16.gguf876 MBalways needed

The "fits" column is weights + 0.86 GiB projector + 0.56 GiB KV cache against 7.91 GiB of VRAM.

This is the line. Nine of the eleven quants run on an 8 GB card; Q6_K and Q8_0 do not, and they are the two that need a 12 GB card. Q5_K_M (7.44 GiB) is the practical ceiling — above that there is no room for anything else on the machine.

Q4_K_M is the default and the one the measurements below were taken on. The vision projector is always f16 — 637 MB, and quantizing it saves little while costing image fidelity.

Why the range matters here: the fp16 checkpoint does not fit on an 8 GB card alongside a KV cache and needs ~10 GB free to load at all. Q3KM is 1.02 GB and runs on a 3 GB card with room for context. For a decision model the readouts are logits at one position, so the lighter quants hold up better than they would for long-form generation — but the calibration below was fitted on the unquantized path, so treat lower quants as untested on it.

Run it

bash
# llama.cpp — needs BOTH files (swap the quant as needed)
llama-server -m imajev-9b-Q5_K_M.gguf \
             --mmproj imajev-9b-mmproj-f16.gguf \
             --host 0.0.0.0 --port 8080 -ngl 99 -c 2048 -np 1

# On Pascal / sm_61 add this, or the vision encoder segfaults (exit 139):
#   --no-mmproj-offload

Ollama 0.35 does not accept an ADAPTER line for an mmproj GGUF ("LoRA adapters are no longer supported"), so a Multimodal GGUF cannot be attached that way. Use llama.cpp for the image path.

Decision readout

Imajev is a decision model: it returns a probability per option instead of generating text. Read the option logits at the last position and divide by the calibration temperature before the softmax:

temperature = 1.748        (calibration_version imajev-9b-1.0,
                            from boolean:2 in calibration.json)

That division happens client-side. It is not stored in the GGUF, and it is what makes the numbers calibrated rather than merely ordered.

Measured on a GTX 1070 (Pascal, sm_61, 8 GB). Q4KM plus the projector is 6.12 GiB of 7.91, with 2.92 GiB free afterwards.

black   A 0.222  B 0.101  C 0.254  D 0.423
white   A 0.180  B 0.127  C 0.223  D 0.470
red     A 0.309  B 0.123  C 0.232  D 0.337
page 1  A 0.496  B 0.180  C 0.188  D 0.136
page 2  A 0.545  B 0.152  C 0.178  D 0.124

Four images, five distributions: the image reaches the forward pass.

Worth knowing next to the 2B and 4B numbers: the 9B separates less sharply. Black scores D 0.42 here against 0.89 on the 2B and 0.73 on the 4B, and the pages score A 0.50 against 0.48. The larger model is the more cautious judge on this hardware, which is worth more than a confident wrong answer — but it also means the separation between "not a page" and "a page" is narrower, so its margin is the weaker signal of the three.

Requires --no-mmproj-offload on sm_61, like the other two.

Calibration

calibration.json (imajev-9b-1.0) ships separately — per question type and per option count, not one global number. The vision modality has its own table upstream; keep the two apart.

Provenance

Built from `mohit67890/imajev-9b` — its own adapter, not a derivative of the 2B or 4B — merged onto Qwen/Qwen3.5-9B, revision c202236, and exported with llama.cpp convert_hf_to_gguf.py.

The three sizes share an upstream decision method but are separate LoRA adapters over different base sizes (24 / 36 / 32 layers). Each was merged into its own base; a 2B adapter cannot produce 9B weights. The full recipe, including the five pitfalls that cost hours, is in build_imajev_gguf.sh in this repository — run it top to bottom before trusting any of it.

The five pitfalls

  1. 1.Look the architecture up in the table, not in a filename. conversion/__init__.py maps Qwen3_5ForConditionalGeneration to qwen. There is no qwen35.py. The converter classes live flat in conversion/, not in conversion/models/.
  2. 2.`--mmproj` writes the projector and drops the weights. For the text model it yields 0.62 GB instead of 3.52 GB. Do not use it there.
  3. 3.Qwen3.5 announces an MTP layer that the checkpoint does not contain. mtp_num_hidden_layers = 1, but zero mtp.* tensors. The converter counts the config (24 + 1 = 25) and exports 24, so the loader goes looking for blk.24.attn_norm.weight. --no-mtp fixes the text path — and that flag is not defined for the vision architecture ("--mtp / --no-nextn are not supported for Qwen3_5ForConditionalGeneration").
  4. 4.Patching `block_count` to 24 by hand creates a second bug. The loader then hunts for blk.23.nextn.*. The original 25 was correct; the model was incomplete, not the number.
  5. 5.`image input is not supported` means the mmproj is missing at server start, not that vision is broken.

Merging note: use AutoModelForImageTextToText, not AutoModelForCausalLM — the adapter targets .*language_model.*.(q_proj|…) and the causal class drops that wrapper. Set architectures to Qwen3_5ForConditionalGeneration so the converter takes the vision path.

License

Apache-2.0, following the upstream adapter. Built on Qwen3.5-2B (Apache-2.0) with the adapter from mohit67890/imajev-2b (Apache-2.0).