Team Ai
Modelpublic

VertexAGI/prism-caption-3-micro

sourceHugging Faceotherupdated 9d agoView on Hugging Face
0likes220downloads
Model Card

Prism Caption 3 Micro

Prism Caption 3 Micro is a chat-titling model: given the first user message of a conversation, it writes a short, specific, correctly formatted title (4-6 words, title case, naming the actual subject). It is the same ~350M model as Prism Caption 2.5 Micro, retrained on a 30,000-example dataset, fine-tuned with LoRA on LFM2-350M. Part of the Prism family of small, single-purpose models.

Evaluation

Prism Caption 3 comes in two sizes, Micro (354M) and Pico (135M). Both are compared below with their untuned base models and the previous generation, on the same inputs: 275 held-out topics that never appear in any training bank, and 30 hand-written, realistic multi-sentence chat openers (the training prompts are short templated phrasings, so the second set checks that the model generalizes beyond the template).

275 held-out topics

SystemFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
Base LFM2-350M231/275247/27590/27510.269 ms
Prism Caption 2.5 Micro (published)27/275274/275222/2755.252 ms
Prism Caption 3 Micro0/275275/275264/2754.963 ms
Base SmolLM2-135M273/275255/2754/27520.8121 ms
Prism Caption 3 Pico2/275275/275219/2755.348 ms

30 realistic multi-sentence openers

SystemFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
Base LFM2-350M25/3027/309/3014.285 ms
Prism Caption 2.5 Micro (published)9/3029/3019/306.860 ms
Prism Caption 3 Micro0/3030/3027/304.864 ms
Base SmolLM2-135M29/3029/301/3020.3126 ms
Prism Caption 3 Pico2/3029/3025/304.747 ms

"Format issues" = formatting problems (too long, too terse, leaked preamble, trailing punctuation, multiline). "Relevant" = the title shares a content word with the message. "3-6 words" = title length within the target range. "Speed" = average wall-clock time to generate one title (up to 28 new tokens, greedy decoding, single request) with mlx-lm on an Apple M4 (16 GB); speeds from separate runs differ by roughly +/-15 ms, so treat gaps smaller than that as noise. The Caption 3 rows above are the published default build (MLX 6-bit); every build is in the table further down.

The relevance check is a word-overlap rule and the format check is rule-based: this is a regression-style check, not a human or model judge, and both Caption 3 sizes are near its ceiling. The Caption 2.5 Micro row was measured on the published (fused, 4-bit) weights; the 0/275 issues figure on its own model card was measured earlier with the unfused adapter loaded.

Formats

This repo holds both an MLX and a GGUF build, plus extra MLX versions:

FormatLocationNotes
MLX 6-bit (default)repo root (model.safetensors + config)mlx-lm on Apple Silicon. Best balance: no measurable quality loss versus fp16 at lower latency
MLX fp16mlx-fp16/Full-precision reference
MLX 4-bit, group size 32mlx-4bit-g32/Fastest MLX option; small quality trade-off versus 6-bit (see the table below)
GGUF Q4_K_Mprism_caption_3_micro_Q4_K_M.ggufllama.cpp and compatible runtimes (LM Studio, Ollama, ...)
GGUF Q8_0prism_caption_3_micro_Q8_0.ggufHigher-fidelity GGUF

Why not plain 4-bit

The model was fine-tuned on a 4-bit base and then fused. Fusing a LoRA into 4-bit weights and re-quantizing with the default settings loses part of the fine-tune: the plain 4-bit MLX build measured 16/275 formatting issues versus 0/275 for fp16 (same model, same prompts), so it is not published. The builds here fuse the LoRA at full precision first and quantize afterwards. The 6-bit build holds quality; the group-32 4-bit build is the faster option and trades a little quality for it (see the table: lengths loosen a bit, and Pico picks up a few format issues). GGUF K-quants held quality too.

All builds measured

275 held-out topics

BuildFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
MLX 6-bit (repo root, default)0/275275/275264/2754.963 ms
MLX fp16 (mlx-fp16/)0/275275/275267/2754.9105 ms
MLX 4-bit group-32 (mlx-4bit-g32/)0/275275/275240/2755.356 ms
GGUF Q8_00/275275/275267/2754.953 ms
GGUF Q4KM0/275275/275236/2755.377 ms
MLX 8-bit (measured, not published)0/275275/275267/2754.869 ms
MLX plain 4-bit (measured, not published)16/275275/275205/2755.654 ms

30 realistic multi-sentence openers

BuildFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
MLX 6-bit (repo root, default)0/3030/3027/304.864 ms
MLX fp16 (mlx-fp16/)1/3030/3027/305.0106 ms
MLX 4-bit group-32 (mlx-4bit-g32/)0/3030/3024/305.357 ms
GGUF Q8_01/3030/3027/305.077 ms
GGUF Q4KM0/3030/3027/305.059 ms
MLX 8-bit (measured, not published)1/3030/3027/304.974 ms
MLX plain 4-bit (measured, not published)8/3030/3020/306.960 ms

GGUF rows were measured through llama-server (Metal), which adds a few ms of local HTTP overhead; MLX rows through mlx-lm.

Usage -- MLX

python
from mlx_lm import load, generate

model, tokenizer = load("VertexAGI/prism-caption-3-micro")   # 6-bit default. For mlx-fp16/ or mlx-4bit-g32/, download the repo and pass that subfolder's local path

messages = [{"role": "system", "content": (
    "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title."
)}, {"role": "user", "content": "Any advice on how to fix a leaking kitchen faucet?"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=text, max_tokens=24))

Usage -- GGUF (llama.cpp)

bash
hf download VertexAGI/prism-caption-3-micro prism_caption_3_micro_Q4_K_M.gguf --local-dir .
llama-cli -m prism_caption_3_micro_Q4_K_M.gguf -st -n 24 --temp 0 \
  -sys "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title." \
  -p "Any advice on how to fix a leaking kitchen faucet?"

Model Details

Base modelLiquidAI/LFM2-350M
Fine-tuning base checkpointthe 4-bit MLX checkpoint mlx-community/LFM2-350M-4bit
ArchitectureLFM2: hybrid short-convolution / attention
Fine-tuning methodLoRA (rank 8, scale 20.0, 16 (full depth) layers; 2.998M (0.846%) trainable parameters)
FrameworkMLX / mlx-lm, on Apple Silicon
LicenseLFM Open License v1.0, inherited from the LFM2-350M base model. Free for research/non-commercial use and for commercial use under $10M annual revenue.

Training Data

A 30,000-example chat-titling dataset (27,000 train / 3,000 validation), the 13,000-example set behind Prism Caption 2.5 Micro extended with 17,000 new examples over a much larger topic bank: 9,207 unique topics (up from 1,207), 28,367 unique prompts, 23,241 unique titles. Each example is a first user message paired with a teacher-written title. Teachers (fast models, cycled/switched adaptively by recent success rate):

TeacherExamplesShare
openai/gpt-oss-20b (NIM)23,69079.0%
nvidia/nemotron-3.5-lightning-30b-a3b (NIM)4,25014.2%
poolside/laguna-s-2.1:free (OpenRouter)1,0273.4%
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning (NIM)6002.0%
openai/gpt-oss-120b (NIM, retired 2026-09-03)4331.4%

New examples were filtered for format (3-8 words, no preamble) and relevance (the title must share a content word with the message). The 275 evaluation topics are excluded from every training bank.

Training Procedure

  • —Method: LoRA, rank 8, scale 20.0, dropout 0.0, 16 (full depth) layers; Adam, learning rate 1e-5, batch size 4, sequence length 256
  • —Steps: 6,750 iterations (exactly one epoch of the 27,000 training examples), validation every 250 steps
  • —Validation loss: 7.122 at initialization, best 0.333 at iteration 6,500 (final 0.345) -- the best checkpoint was used
  • —Throughput: ~2.35 it/s, ~1,050 tokens/s, peak memory ~1.2 GB (Apple M4)
  • —Hyperparameters are identical to Prism Caption 2.5 Micro, so the comparison with it isolates the data (13k to 30k examples, larger topic bank) rather than the recipe.
  • —Release builds: the LoRA adapter was fused into the de-quantized base at full precision, then quantized (MLX 6-bit / group-32 4-bit) or converted to GGUF (Q80 / Q4K_M) from that fp16 model.

Eval scripts and raw results for every build are in eval/.

Limitations

Trained on synthetic titles distilled from a mix of teacher models, so some stylistic inconsistency between teachers may remain. Validated only on English, conversational, everyday-topic inputs; highly technical or non-English inputs are untested. Titles are short and extractive-leaning; a message with several unrelated requests will get a title for only one of them. Evaluation uses rule-based checks on 275 + 30 prompts, which are near ceiling for both sizes, so differences between Micro and Pico on this eval are small and should not be over-read.

License

LFM Open License v1.0, inherited from the LFM2-350M base model. Free for research/non-commercial use and for commercial use under $10M annual revenue.