Team Ai
Modelpublic

VertexAGI/prism-caption-3-pico

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes244downloads
Model Card

Prism Caption 3 Pico

Prism Caption 3 Pico is a chat-titling model: given the first user message of a conversation, it writes a short, specific, correctly formatted title (4-6 words, title case, naming the actual subject). It is an experimental 135M-parameter sibling, less than 40% the size of Micro, fine-tuned with LoRA on SmolLM2-135M-Instruct. Part of the Prism family of small, single-purpose models.

Evaluation

Prism Caption 3 comes in two sizes, Micro (354M) and Pico (135M). Both are compared below with their untuned base models and the previous generation, on the same inputs: 275 held-out topics that never appear in any training bank, and 30 hand-written, realistic multi-sentence chat openers (the training prompts are short templated phrasings, so the second set checks that the model generalizes beyond the template).

275 held-out topics

SystemFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
Base LFM2-350M231/275247/27590/27510.269 ms
Prism Caption 2.5 Micro (published)27/275274/275222/2755.252 ms
Prism Caption 3 Micro0/275275/275264/2754.963 ms
Base SmolLM2-135M273/275255/2754/27520.8121 ms
Prism Caption 3 Pico2/275275/275219/2755.348 ms

30 realistic multi-sentence openers

SystemFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
Base LFM2-350M25/3027/309/3014.285 ms
Prism Caption 2.5 Micro (published)9/3029/3019/306.860 ms
Prism Caption 3 Micro0/3030/3027/304.864 ms
Base SmolLM2-135M29/3029/301/3020.3126 ms
Prism Caption 3 Pico2/3029/3025/304.747 ms

"Format issues" = formatting problems (too long, too terse, leaked preamble, trailing punctuation, multiline). "Relevant" = the title shares a content word with the message. "3-6 words" = title length within the target range. "Speed" = average wall-clock time to generate one title (up to 28 new tokens, greedy decoding, single request) with mlx-lm on an Apple M4 (16 GB); speeds from separate runs differ by roughly +/-15 ms, so treat gaps smaller than that as noise. The Caption 3 rows above are the published default build (MLX 6-bit); every build is in the table further down.

The relevance check is a word-overlap rule and the format check is rule-based: this is a regression-style check, not a human or model judge, and both Caption 3 sizes are near its ceiling. The Caption 2.5 Micro row was measured on the published (fused, 4-bit) weights; the 0/275 issues figure on its own model card was measured earlier with the unfused adapter loaded.

Formats

This repo holds both an MLX and a GGUF build, plus extra MLX versions:

FormatLocationNotes
MLX 6-bit (default)repo root (model.safetensors + config)mlx-lm on Apple Silicon. Best balance: no measurable quality loss versus fp16 at lower latency
MLX fp16mlx-fp16/Full-precision reference
MLX 4-bit, group size 32mlx-4bit-g32/Fastest MLX option; small quality trade-off versus 6-bit (see the table below)
GGUF Q4_K_Mprism_caption_3_pico_Q4_K_M.ggufllama.cpp and compatible runtimes (LM Studio, Ollama, ...)
GGUF Q8_0prism_caption_3_pico_Q8_0.ggufHigher-fidelity GGUF

Why not plain 4-bit

The model was fine-tuned on a 4-bit base and then fused. Fusing a LoRA into 4-bit weights and re-quantizing with the default settings loses part of the fine-tune: the plain 4-bit MLX build measured 15/275 formatting issues versus 1/275 for fp16 (same model, same prompts), so it is not published. The builds here fuse the LoRA at full precision first and quantize afterwards. The 6-bit build holds quality; the group-32 4-bit build is the faster option and trades a little quality for it (see the table: lengths loosen a bit, and Pico picks up a few format issues). GGUF K-quants held quality too.

All builds measured

275 held-out topics

BuildFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
MLX 6-bit (repo root, default)2/275275/275219/2755.348 ms
MLX fp16 (mlx-fp16/)1/275275/275221/2755.461 ms
MLX 4-bit group-32 (mlx-4bit-g32/)7/275274/275216/2755.547 ms
GGUF Q8_01/275275/275225/2755.361 ms
GGUF Q4KM1/275275/275230/2755.359 ms
MLX 8-bit (measured, not published)2/275275/275226/2755.350 ms
MLX plain 4-bit (measured, not published)15/275275/275194/2755.846 ms

30 realistic multi-sentence openers

BuildFormat issuesRelevant3-6 wordsAvg wordsSpeed (per title)
MLX 6-bit (repo root, default)2/3029/3025/304.747 ms
MLX fp16 (mlx-fp16/)1/3029/3027/304.860 ms
MLX 4-bit group-32 (mlx-4bit-g32/)0/3029/3024/305.749 ms
GGUF Q8_01/3029/3027/304.553 ms
GGUF Q4KM1/3029/3027/304.751 ms
MLX 8-bit (measured, not published)1/3029/3026/304.750 ms
MLX plain 4-bit (measured, not published)5/3029/3021/305.747 ms

GGUF rows were measured through llama-server (Metal), which adds a few ms of local HTTP overhead; MLX rows through mlx-lm.

Usage -- MLX

python
from mlx_lm import load, generate

model, tokenizer = load("VertexAGI/prism-caption-3-pico")   # 6-bit default. For mlx-fp16/ or mlx-4bit-g32/, download the repo and pass that subfolder's local path

messages = [{"role": "system", "content": (
    "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title."
)}, {"role": "user", "content": "Any advice on how to fix a leaking kitchen faucet?"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=text, max_tokens=24))

Usage -- GGUF (llama.cpp)

bash
hf download VertexAGI/prism-caption-3-pico prism_caption_3_pico_Q4_K_M.gguf --local-dir .
llama-cli -m prism_caption_3_pico_Q4_K_M.gguf -st -n 24 --temp 0 \
  -sys "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title." \
  -p "Any advice on how to fix a leaking kitchen faucet?"

Model Details

Base modelHuggingFaceTB/SmolLM2-135M-Instruct
Fine-tuning base checkpointthe bf16 MLX checkpoint mlx-community/SmolLM2-135M-Instruct
ArchitectureSmolLM2: Llama-style decoder, 30 layers
Fine-tuning methodLoRA (rank 8, scale 20.0, 30 (full depth) layers; 2.442M (1.816%) trainable parameters)
FrameworkMLX / mlx-lm, on Apple Silicon
LicenseApache-2.0, inherited from SmolLM2-135M-Instruct.

Training Data

A 30,000-example chat-titling dataset (27,000 train / 3,000 validation), the 13,000-example set behind Prism Caption 2.5 Micro extended with 17,000 new examples over a much larger topic bank: 9,207 unique topics (up from 1,207), 28,367 unique prompts, 23,241 unique titles. Each example is a first user message paired with a teacher-written title. Teachers (fast models, cycled/switched adaptively by recent success rate):

TeacherExamplesShare
openai/gpt-oss-20b (NIM)23,69079.0%
nvidia/nemotron-3.5-lightning-30b-a3b (NIM)4,25014.2%
poolside/laguna-s-2.1:free (OpenRouter)1,0273.4%
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning (NIM)6002.0%
openai/gpt-oss-120b (NIM, retired 2026-09-03)4331.4%

New examples were filtered for format (3-8 words, no preamble) and relevance (the title must share a content word with the message). The 275 evaluation topics are excluded from every training bank.

Training Procedure

  • —Method: LoRA, rank 8, scale 20.0, dropout 0.0, 30 (full depth) layers; Adam, learning rate 1e-5, batch size 4, sequence length 256
  • —Steps: 6,750 iterations (exactly one epoch of the 27,000 training examples), validation every 250 steps
  • —Validation loss: 3.873 at initialization, best 0.364 at iteration 5,000 (final 0.378) -- the best checkpoint was used
  • —Throughput: ~4.5 it/s, ~1,960 tokens/s, peak memory ~0.9 GB (Apple M4)
  • —Hyperparameters are identical to Prism Caption 2.5 Micro, so the comparison with it isolates the data (13k to 30k examples, larger topic bank) rather than the recipe.
  • —Release builds: the LoRA adapter was fused into the de-quantized base at full precision, then quantized (MLX 6-bit / group-32 4-bit) or converted to GGUF (Q80 / Q4K_M) from that fp16 model.

Eval scripts and raw results for every build are in eval/.

Limitations

Trained on synthetic titles distilled from a mix of teacher models, so some stylistic inconsistency between teachers may remain. Validated only on English, conversational, everyday-topic inputs; highly technical or non-English inputs are untested. Titles are short and extractive-leaning; a message with several unrelated requests will get a title for only one of them. Evaluation uses rule-based checks on 275 + 30 prompts, which are near ceiling for both sizes, so differences between Micro and Pico on this eval are small and should not be over-read. This is the experimental, smallest model in the Prism Caption line.

License

Apache-2.0, inherited from SmolLM2-135M-Instruct.