VertexAGI/prism-caption-3-pico
Prism Caption 3 Pico
Prism Caption 3 Pico is a chat-titling model: given the first user message of a conversation, it writes a short, specific, correctly formatted title (4-6 words, title case, naming the actual subject). It is an experimental 135M-parameter sibling, less than 40% the size of Micro, fine-tuned with LoRA on SmolLM2-135M-Instruct. Part of the Prism family of small, single-purpose models.
Evaluation
Prism Caption 3 comes in two sizes, Micro (354M) and Pico (135M). Both are compared below with their untuned base models and the previous generation, on the same inputs: 275 held-out topics that never appear in any training bank, and 30 hand-written, realistic multi-sentence chat openers (the training prompts are short templated phrasings, so the second set checks that the model generalizes beyond the template).
275 held-out topics
30 realistic multi-sentence openers
"Format issues" = formatting problems (too long, too terse, leaked preamble, trailing punctuation, multiline). "Relevant" = the title shares a content word with the message. "3-6 words" = title length within the target range. "Speed" = average wall-clock time to generate one title (up to 28 new tokens, greedy decoding, single request) with mlx-lm on an Apple M4 (16 GB); speeds from separate runs differ by roughly +/-15 ms, so treat gaps smaller than that as noise. The Caption 3 rows above are the published default build (MLX 6-bit); every build is in the table further down.
The relevance check is a word-overlap rule and the format check is rule-based: this is a regression-style check, not a human or model judge, and both Caption 3 sizes are near its ceiling. The Caption 2.5 Micro row was measured on the published (fused, 4-bit) weights; the 0/275 issues figure on its own model card was measured earlier with the unfused adapter loaded.
Formats
This repo holds both an MLX and a GGUF build, plus extra MLX versions:
Why not plain 4-bit
The model was fine-tuned on a 4-bit base and then fused. Fusing a LoRA into 4-bit weights and re-quantizing with the default settings loses part of the fine-tune: the plain 4-bit MLX build measured 15/275 formatting issues versus 1/275 for fp16 (same model, same prompts), so it is not published. The builds here fuse the LoRA at full precision first and quantize afterwards. The 6-bit build holds quality; the group-32 4-bit build is the faster option and trades a little quality for it (see the table: lengths loosen a bit, and Pico picks up a few format issues). GGUF K-quants held quality too.
All builds measured
275 held-out topics
30 realistic multi-sentence openers
GGUF rows were measured through llama-server (Metal), which adds a few ms of local HTTP overhead; MLX rows through mlx-lm.
Usage -- MLX
from mlx_lm import load, generate
model, tokenizer = load("VertexAGI/prism-caption-3-pico") # 6-bit default. For mlx-fp16/ or mlx-4bit-g32/, download the repo and pass that subfolder's local path
messages = [{"role": "system", "content": (
"You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title."
)}, {"role": "user", "content": "Any advice on how to fix a leaking kitchen faucet?"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=text, max_tokens=24))Usage -- GGUF (llama.cpp)
hf download VertexAGI/prism-caption-3-pico prism_caption_3_pico_Q4_K_M.gguf --local-dir .
llama-cli -m prism_caption_3_pico_Q4_K_M.gguf -st -n 24 --temp 0 \
-sys "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title." \
-p "Any advice on how to fix a leaking kitchen faucet?"Model Details
Training Data
A 30,000-example chat-titling dataset (27,000 train / 3,000 validation), the 13,000-example set behind Prism Caption 2.5 Micro extended with 17,000 new examples over a much larger topic bank: 9,207 unique topics (up from 1,207), 28,367 unique prompts, 23,241 unique titles. Each example is a first user message paired with a teacher-written title. Teachers (fast models, cycled/switched adaptively by recent success rate):
New examples were filtered for format (3-8 words, no preamble) and relevance (the title must share a content word with the message). The 275 evaluation topics are excluded from every training bank.
Training Procedure
- Method: LoRA, rank 8, scale 20.0, dropout 0.0, 30 (full depth) layers; Adam, learning rate 1e-5, batch size 4, sequence length 256
- Steps: 6,750 iterations (exactly one epoch of the 27,000 training examples), validation every 250 steps
- Validation loss: 3.873 at initialization, best 0.364 at iteration 5,000 (final 0.378) -- the best checkpoint was used
- Throughput: ~4.5 it/s, ~1,960 tokens/s, peak memory ~0.9 GB (Apple M4)
- Hyperparameters are identical to Prism Caption 2.5 Micro, so the comparison with it isolates the data (13k to 30k examples, larger topic bank) rather than the recipe.
- Release builds: the LoRA adapter was fused into the de-quantized base at full precision, then quantized (MLX 6-bit / group-32 4-bit) or converted to GGUF (Q80 / Q4K_M) from that fp16 model.
Eval scripts and raw results for every build are in eval/.
Limitations
Trained on synthetic titles distilled from a mix of teacher models, so some stylistic inconsistency between teachers may remain. Validated only on English, conversational, everyday-topic inputs; highly technical or non-English inputs are untested. Titles are short and extractive-leaning; a message with several unrelated requests will get a title for only one of them. Evaluation uses rule-based checks on 275 + 30 prompts, which are near ceiling for both sizes, so differences between Micro and Pico on this eval are small and should not be over-read. This is the experimental, smallest model in the Prism Caption line.
License
Apache-2.0, inherited from SmolLM2-135M-Instruct.
