Team Ai
Modelpublic

dmis-lab/Gemma4-12B-CVRR

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes1kdownloads
Model Card

Gemma4-12B-CVRR

<p align="center"> <a href="https://arxiv.org/abs/2609.06746"><strong>๐Ÿ“„ Paper</strong></a> ยท <a href="https://github.com/dmis-lab/CVRR"><strong>๐Ÿ’ป Code</strong></a> </p>

Introduction

Latent visual reasoning is not just about storing visual information in hidden states. That information should matter for the model's answer. When an answer decoder can access the original multimodal context through a parallel path, the presence of informative latent states alone does not establish their use.

Our work, Reason Through the Latent! Making Latent Visual Reasoning Necessary, introduces Causal Visual Recurrent Reasoning (CVRR). CVRR separates visual access during reasoning from visual access during answering, focusing on three design principles:

๐Ÿ‘๏ธ Preserving pretrained visual competence Reasoning starts from the native image-conditioned question representation, preserving visual information already integrated by the pretrained VLM.

๐Ÿ” Refining reasoning with persistent visual evidence A shared native decoder layer updates the question state while retaining access to fixed visual evidence. The recurrent transition is adapted with LoRA and answer-token supervision.

๐Ÿ”’ Making the reasoning state the visual answer interface The upper answer decoder receives the final recurrent question state, without direct access to the original visual rows or original multimodal prefix caches. The lower-layer prefix context used for answer generation follows a text-only path.

Model Checkpoints

Final Models

We release CVRR checkpoints based on the following pretrained vision-language backbones. Each repository contains the native backbone weights, the merged recurrent transition, and the custom inference code; no separate LoRA download is required.

This checkpoint is based on `google/gemma-4-12B-it`.

Loading

Install the model-specific dependencies from requirements.txt, then load directly from the Hub:

bash
pip install torch==2.9.1 transformers==5.10.4 peft timm einops accelerate safetensors huggingface_hub Pillow numpy sentencepiece

The first from_pretrained call downloads and caches the complete model package.

python
import torch
from PIL import Image
from transformers import AutoModelForImageTextToText

model = AutoModelForImageTextToText.from_pretrained(
    "dmis-lab/Gemma4-12B-CVRR",
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="cuda:0",
).eval()

image = Image.open("example.jpg").convert("RGB")
inputs = model.prepare_inputs(
    image,
    "What is the dominant color? A. Red B. Blue C. Green D. Yellow. Answer with the option letter.",
)
answer_ids = model.generate(
    **inputs,
    do_sample=False,
    max_new_tokens=64,
    eos_token_id=model.tokenizer.eos_token_id,
)
print(model.tokenizer.decode(answer_ids[0], skip_special_tokens=True))

This model supports strict first-answer-token readout and deterministic greedy answer generation. The default recurrent update coefficient is beta=0.5.

Use one complete model replica per GPU. This custom Transformers loader does not support automatic model sharding, generic save_pretrained() reserialization, or direct loading through vLLM/SGLang. Pin the Hub revision for reproducible use.

License

Please follow the upstream model's terms and the included `NOTICE.txt` and `licenses/` attribution files. The component licenses remain applicable to their respective materials.

Citation

bibtex
@misc{park2026reasonlatentmakinglatent,
      title={Reason Through the Latent! Making Latent Visual Reasoning Necessary},
      author={Suhyeong Park and Junha Jung and Jaewoo Kang},
      year={2026},
      eprint={2609.06746},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.06746},
}