textilelabs/Loom-Spark-3.2
<div align="center"> <img src="banner.jpg" alt="Loom Spark 3.2" width="520"> </div>
Loom Spark 3.2
<img src="logo.jpg" alt="" width="20" height="20" style="border-radius:4px;vertical-align:middle;margin-right:6px;"> 22.8M parameters · 20 layers · 2,048 context · Textile Labs
The Spark that knows what it doesn't know. Successor to Loom Spark 3. Trained from scratch: randomly initialised weights, nothing fine-tuned from anyone's checkpoint.
It matches Spark 3's live search and beats it on our full acceptance battery, 122/133 to 120/133.
What changed from Spark 3
Spark 3 only declined capital-city questions offline; for everything else it guessed. Spark 3.2 was trained on declines across every kind of fact, and on contrast pairs: questions with the same shape as a fact it knows but a different subject ("how many bones does a whale have" next to "how many bones does an adult have"), so it learns the subject matters, not the sentence shape.
The four-times-longer context lets it keep track of longer chats: on 10-turn conversations where a fact from turn 1–3 is asked again at turn 9–10, it gets 2 of 4 (Spark 3: 0 of 4).
The search harness
The model decides a search is needed and writes the query. harness.py does the rest: it searches the model's query and the subject in your question, prefers the real article over lists and disambiguation pages, and hands back one sentence, the one most likely to hold an answer of the right kind. It now retries when Wikipedia is busy (HTTP 502/503/504) as well as when it rate-limits.
you who composed the four seasons
Loom Spark 3.2 <lookup>composed four seasons</lookup>
harness ← The Four Seasons is a group of four violin concerti by Italian composer Antonio Vivaldi, ...
Loom Spark 3.2 Antonio Vivaldi. I looked that one up.Measured against Spark 3
Same tests, same harness, same settings, both models through Ollama, 2026-10-01. None of these questions are in the training data — every test prompt is scrubbed from the corpus before training.
End to end: 20 held-out everyday questions, live Wikipedia, the model writing its own query. Scored on the final answer.
Read by eye, one of Spark 3.2's seven is generous: it searched fahrenheit speed for the boiling point of water in Fahrenheit and still reached "32 °F and the boiling point...".
The acceptance battery, row by row:
Held-out behaviour tests, written before Spark 3.2 was trained:
Spark 3's 20/20 on the same-shape row is mostly empty: asked "how far is mars from earth", it loops ("mars, mars, mars…") rather than claiming anything. Spark 3.2 declines most of these properly; its one miss is below.
Every Loom text model
† scored on a battery with every test prompt scrubbed from training. Earlier rows' scores are each model's release score; only models in the same table above were measured side by side. Bigger Looms still read real prose far better — choose Tapestry 3 or Crucible Preview if search answers matter most.
Read this before you use it
Every point here was measured.
- It gets about a third of everyday questions right with search. Same as Spark 3. "I looked that up" means it searched, not that it read the result correctly. Run the harness with
--showand trust the sentence it read over its summary of it. - It reads the right sentence and picks the wrong part. Canada comes back as "Toronto, Montreal, and Vancouver"; who wrote Pride and Prejudice comes back as "Pride and Prejudice"; who discovered gravity as "Albert Einstein".
- Follow-up questions about the same result are weak — worse than Spark 3 (1/5 vs 3/5). Ask a fresh, complete question instead of "how many people live there".
- It usually doesn't say when a result lacks the answer (2/5). It answers from whatever it read.
- Long pasted documents don't work yet. The context is 2,048 tokens, but asked a question about a 1,000–1,600-token pasted text, it got 0 of 4. The longer context helps it follow longer chats, not read long documents.
- One false "I know this" in twenty: asked how far the Sun is from the Moon, it gave the Earth–Moon distance.
- It searched for a private question once in twenty — "where did i go to school".
- Warmth is a coin flip. Half the time good news gets "Okay, I'll remember that." instead of congratulations.
- It sometimes garbles a query —
calyfor the capital of italy. The harness's subject search catches most of these. - Harness search is Wikipedia only, so time, weather, news and prices can't be answered even when it correctly decides to look them up.
Usage — the harness
python3 harness.py "whats the capital of peru"
python3 harness.py # interactive
python3 harness.py --show "who wrote hamlet" # see what it searched and read
python3 harness.py --no-tools "who are you"Stdlib only. Wikipedia needs no API key. Swap search() for anything — the contract is text in, one sentence out. Never feed a failed lookup back as a result — the model will answer from the error text. harness.py fails loudly instead.
Usage — Ollama
ollama run hf.co/textilelabs/Loom-Spark-3.2 "who are you"template and params are read automatically. Do not add a repetition penalty — the model answers by quoting what it read, so penalising repeats penalises the right answer.
Usage — transformers
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("textilelabs/Loom-Spark-3.2")
model = AutoModelForCausalLM.from_pretrained("textilelabs/Loom-Spark-3.2").eval()
eot = tok.convert_tokens_to_ids("<|eot|>")
def ask(message, tools=False):
p = f"<tools:{'on' if tools else 'off'}>\n<user>\n{message}\n<|eot|>\n<loom>\n"
ids = tok(p, return_tensors="pt", add_special_tokens=False).input_ids
with torch.no_grad():
out = model.generate(ids, max_new_tokens=96, do_sample=False, eos_token_id=eot,
pad_token_id=tok.convert_tokens_to_ids("<|pad|>"))[0]
return tok.decode(out[ids.shape[1]:], skip_special_tokens=False).replace("<|eot|>", "").strip()Prompt format is exact: <tools:off>\n<user>\n{message}\n<|eot|>\n<loom>\n. For more turns, append {reply}<|eot|>\n<user>\n{next message}\n<|eot|>\n<loom>\n.
How it was built
The per-conversation mask mattered more than anything else. Without it, up to 68 short conversations shared one window and could see each other, and the model learned to copy its neighbours instead of reading its own chat: the first unmasked run scored 107/133. Same data with the mask: 120/133.
Files
config.json / model.safetensors the model
tokenizer.json / tokenizer_config.json custom BPE tokenizer, 4,096 tokens
loom-spark-3.2-f16.gguf for Ollama / llama.cpp
harness.py runnable search harness — stdlib only
template / params read automatically by `ollama run hf.co/...`
Modelfile for building locally
ATTRIBUTION.md required credits for the training corporaTraining data
Every search query is derived mechanically from these sources. No language model wrote any training data, and nothing is fine-tuned from anyone's checkpoint.
License
Model: MIT. Training data retains its original licences and attribution.
