mkd-hika/keural-vision-encoder-poc
<div align="center">
๐๏ธ Keural Vision Encoder <sub>(PoC ยท V0.1)</sub>
A 24.7M-parameter vision encoder trained _from scratch_ โ no pretrained backbone, no CLIP weights. Structured tokens that plug into any LLM.
<p> <img src="https://img.shields.io/badge/MKDCo.,Ltd.-Keural-6d28d9?style=for-the-badge" alt="MKD"> <img src="https://img.shields.io/badge/status-V0.1_PoC-16a34a?style=for-the-badge" alt="Status"> </p>
<p> <img src="https://img.shields.io/badge/params-24.7M-8b5cf6" alt="Params"> <img src="https://img.shields.io/badge/from-scratch-ec4899" alt="From scratch"> <img src="https://img.shields.io/badge/embeddim-384-f59e0b" alt="Dim"> <img src="https://img.shields.io/badge/tokenbudget-adaptive-22d3ee" alt="ATB"> <img src="https://img.shields.io/badge/PyTorch-2.x-ee4c2c?logo=pytorch&logoColor=white" alt="PyTorch"> <img src="https://img.shields.io/badge/trainedon-1รRTX_5090-38bdf8" alt="Hardware"> </p>
<br>
<img src="assets/keuralvisionencoder.gif" alt="Where the Keural encoder sits in the VLM pipeline" width="98%">
<br>
The encoder (CNN Stem โ ATB โ Spatial Transformer) feeds the LevelAware Projector โ Mistral-7B in the full [Keural VLM](https://huggingface.co/mkd-hika/keural-vlm-poc).
</div>
โจ Key Innovations
๐ฏ Adaptive Token Budget (ATB) Tokenization โ token count is a runtime parameter. Tokens are allocated to information-dense regions: a blank wall gets fewer, a dense document gets more.
out = encoder(image, token_budget=64) # cheap
out = encoder(image, token_budget=256) # default
out = encoder(image, token_budget=1024) # full fidelity๐ช Hierarchical Concept Tokenization (HCT) โ every token carries a semantic level tag: global (whole-image), region (object-scale), or detail (fine-grained).
out = encoder(image)
print(out.level_ids) # tensor of {0=global, 1=region, 2=detail}โ๏ธ Model Specifications
๐จ Saliency & Token Placement
The ATB tokenizer concentrates tokens on salient regions. Left โ right: original ยท saliency heatmap ยท token placement (global / region / detail).
<div align="center"> <img src="saliencyviz/step36k/cat_saliency.png" alt="Saliency and token placement" width="88%"> </div>
๐งช Training
Phase 1 โ Vision Encoder Pretraining โ
COMPLETE. Trained from scratch for ~75,000 steps on CC3M + CC12M (~6.9M imageโtext pairs), SigLIP-style contrastive objective, 1ร RTX 5090. The frozen encoder is then integrated into the full VLM via LevelAwareProjector (384 โ 2048 โ 4096) โ Mistral-7B-Instruct-v0.3 (4-bit NF4 QLoRA), fine-tuned with SFT (LLaVA-Instruct-150K, 30K steps).
<div align="center"> <img src="assets/traininglosscurves.png" alt="Training loss curves" width="88%"> </div>
๐ Benchmark Results โ SFT-30K
Downstream VLM benchmarks using this encoder (frozen) + projector + Mistral-7B at the SFT-30K checkpoint (supervised fine-tuning, 30,000 steps; before DPO). Evaluated on 1,000 samples each where applicable. The Keural encoder is 12.4ร smaller than LLaVA's CLIP encoder (307M).
These are the SFT-30K numbers (measured from outputs/eval/sft_30k/). Applying DPO alignment (RLHF-V, 3K steps) improves every metric โ e.g. VQAv2 โ 43.6%, ScienceQA โ 53.7%, MME โ 838.8. Full SFT+DPO results, the comparison chart, and the complete VLM live on [mkd-hika/keural-vlm-poc](https://huggingface.co/mkd-hika/keural-vlm-poc).โก Adaptive Token Budget โ latency vs budget
Encode latency vs token budget โ trade visual detail for speed at runtime, no retraining.
<div align="center"> <img src="assets/tokenbudgetcurve.png" alt="ATB token budget vs latency" width="80%"> </div>
๐ Output Contract
Every forward pass returns a KeuralEncoderOutput:
๐ Usage
import torch
from architecture.encoder import KeuralVisionEncoder
from keural_config import KeuralConfig
cfg = KeuralConfig.from_yaml("configs/keural_tiny_poc.yaml")
model = KeuralVisionEncoder.from_pretrained("mkd-hika/keural-vision-encoder-poc")
model.eval()
from PIL import Image
from torchvision import transforms
transform = transforms.Compose([
transforms.Resize((256, 256)),
transforms.ToTensor(),
transforms.Normalize([0.5], [0.5]),
])
image = Image.open("image.jpg").convert("RGB")
x = transform(image).unsqueeze(0) # (1, 3, 256, 256)
with torch.no_grad():
out = model(x)
print(out.tokens.shape) # (1, 256, 384)
print(out.level_ids[0, :8]) # [0, 0, 1, 1, 1, 2, 2, 2]
print(out.pooled.shape) # (1, 384)๐บ๏ธ Roadmap
The mid-level model will use knowledge distillation from SigLIP-400M (natural images) and InternViT-300M (documents/OCR). Teachers are discarded after training โ not part of the final model.
๐ Citation
@misc{keural_vision_encoder_2026,
title = {Keural Vision Encoder: Content-Adaptive Vision Encoding via Saliency-Guided Token Budgets},
author = {Barki, Hika and MKD Co., Ltd.},
year = {2026},
note = {V0.1 โ Phase 1 complete},
}๐ License
Code: MIT. Model weights trained on CC3M + CC12M โ data licenses apply.
<div align="center">
MKD Co., Ltd. โ 2026
</div>
