Team Ai
Modelpublic

textilelabs/Loom-Spark-3.2

sourceHugging Facemitupdated 9d agoView on Hugging Face
0likes608downloads
Model Card

<div align="center"> <img src="banner.jpg" alt="Loom Spark 3.2" width="520"> </div>

Loom Spark 3.2

<img src="logo.jpg" alt="" width="20" height="20" style="border-radius:4px;vertical-align:middle;margin-right:6px;"> 22.8M parameters · 20 layers · 2,048 context · Textile Labs

The Spark that knows what it doesn't know. Successor to Loom Spark 3. Trained from scratch: randomly initialised weights, nothing fine-tuned from anyone's checkpoint.

It matches Spark 3's live search and beats it on our full acceptance battery, 122/133 to 120/133.

What changed from Spark 3

Spark 3**Spark 3.2**
size12.2M22.8M (the Spark tier is now ~20M)
context512 tokens2,048 tokens
hardwareone 2013 desktop CPU, 3 h 54 minKaggle, 2 × T4 GPU, 2 h 13 min
unknown facts, tools offdeclined 3 of 20declines 16 of 20
basic facts it should know0 of 20 right15 of 20 right
talking about itself in its own words9 of 1612 of 16
responding to good and bad news2 of 2010 of 20
prompt injection resisted33 of 3636 of 36

Spark 3 only declined capital-city questions offline; for everything else it guessed. Spark 3.2 was trained on declines across every kind of fact, and on contrast pairs: questions with the same shape as a fact it knows but a different subject ("how many bones does a whale have" next to "how many bones does an adult have"), so it learns the subject matters, not the sentence shape.

The four-times-longer context lets it keep track of longer chats: on 10-turn conversations where a fact from turn 1–3 is asked again at turn 9–10, it gets 2 of 4 (Spark 3: 0 of 4).

The search harness

The model decides a search is needed and writes the query. harness.py does the rest: it searches the model's query and the subject in your question, prefers the real article over lists and disambiguation pages, and hands back one sentence, the one most likely to hold an answer of the right kind. It now retries when Wikipedia is busy (HTTP 502/503/504) as well as when it rate-limits.

you               who composed the four seasons
Loom Spark 3.2    <lookup>composed four seasons</lookup>
harness           ← The Four Seasons is a group of four violin concerti by Italian composer Antonio Vivaldi, ...
Loom Spark 3.2    Antonio Vivaldi. I looked that one up.

Measured against Spark 3

Same tests, same harness, same settings, both models through Ollama, 2026-10-01. None of these questions are in the training data — every test prompt is scrubbed from the corpus before training.

End to end: 20 held-out everyday questions, live Wikipedia, the model writing its own query. Scored on the final answer.

decided to searchwrote its own queryanswer reached the model**answered right**
Spark 320/2020/2012/207/20
Loom Spark 3.220/2020/2013/207/20

Read by eye, one of Spark 3.2's seven is generous: it searched fahrenheit speed for the boiling point of water in Fahrenheit and still reached "32 °F and the boiling point...".

The acceptance battery, row by row:

rowSpark 3**Loom Spark 3.2**
A · says its own name11/1212/12
B · its own name under rough typing (WHATS UR NAME???)11/1211/12
C · 5-turn conversation stays on thread5/55/5
D · answers from a search result3/53/5
E · follow-up answered from the same result3/51/5
F · says it looked, after a lookup5/55/5
G · never claims a lookup it didn't make16/1616/16
H · admits what it can't know about you8/88/8
I · says when a result doesn't contain the answer0/52/5
J · never leaks a search tag with tools off28/2828/28
K · stops on its own12/1212/12
L · searches when it should, not for your private things18/2019/20
total120/133122/133

Held-out behaviour tests, written before Spark 3.2 was trained:

Spark 3**Loom Spark 3.2**
unknown facts, tools off — declines instead of guessing3/2016/20
ten basic facts, tools off — answers right0/2015/20
same facts, tools on — looks them up20/2020/20
same-shape questions it doesn't know — no false "I know this"20/2019/20
a <tools:on> typed inside a message doesn't switch search on12/1212/12
in its own words about itself9/1612/16
warmth — good news and bad news met correctly2/2010/20
prompt injection — kept its identity, didn't obey (12 prompts × 3)33/3636/36
10- and 12-turn conversations — turns answered on target41/4442/44

Spark 3's 20/20 on the same-shape row is mostly empty: asked "how far is mars from earth", it loops ("mars, mars, mars…") rather than claiming anything. Spark 3.2 declines most of these properly; its one miss is below.

Every Loom text model

modelparamsbattery /133live search (e2e)reads real prosestatus
Loom Spark 219.9M~97/1332/20noshipped
Loom Tapestry 222.8M107/133—curated onlyshipped
Loom Tapestry 3 Flash7.18M112/1333/20curated onlyshipped
Loom Spark 3 Flash7.18M119/1335/20curated onlyshipped
Loom Spark 312.2M120/1337/20curated onlyshipped
Loom Weave 259.65M— (method failure)—noshipped (superseded)
Loom Weave 331.5M120/1336/20 heldyes — firstshipped
Loom Tapestry 369.2M123/13312/20 heldyes + multi-hopshipped
Loom Crucible Preview155.0M125/133†10/20 heldyes — best readershipped
Loom Spark 3.222.8M122/133†7/20 heldcurated onlythis model

† scored on a battery with every test prompt scrubbed from training. Earlier rows' scores are each model's release score; only models in the same table above were measured side by side. Bigger Looms still read real prose far better — choose Tapestry 3 or Crucible Preview if search answers matter most.

Read this before you use it

Every point here was measured.

  • —It gets about a third of everyday questions right with search. Same as Spark 3. "I looked that up" means it searched, not that it read the result correctly. Run the harness with --show and trust the sentence it read over its summary of it.
  • —It reads the right sentence and picks the wrong part. Canada comes back as "Toronto, Montreal, and Vancouver"; who wrote Pride and Prejudice comes back as "Pride and Prejudice"; who discovered gravity as "Albert Einstein".
  • —Follow-up questions about the same result are weak — worse than Spark 3 (1/5 vs 3/5). Ask a fresh, complete question instead of "how many people live there".
  • —It usually doesn't say when a result lacks the answer (2/5). It answers from whatever it read.
  • —Long pasted documents don't work yet. The context is 2,048 tokens, but asked a question about a 1,000–1,600-token pasted text, it got 0 of 4. The longer context helps it follow longer chats, not read long documents.
  • —One false "I know this" in twenty: asked how far the Sun is from the Moon, it gave the Earth–Moon distance.
  • —It searched for a private question once in twenty — "where did i go to school".
  • —Warmth is a coin flip. Half the time good news gets "Okay, I'll remember that." instead of congratulations.
  • —It sometimes garbles a query — caly for the capital of italy. The harness's subject search catches most of these.
  • —Harness search is Wikipedia only, so time, weather, news and prices can't be answered even when it correctly decides to look them up.

Usage — the harness

bash
python3 harness.py "whats the capital of peru"
python3 harness.py                              # interactive
python3 harness.py --show "who wrote hamlet"    # see what it searched and read
python3 harness.py --no-tools "who are you"

Stdlib only. Wikipedia needs no API key. Swap search() for anything — the contract is text in, one sentence out. Never feed a failed lookup back as a result — the model will answer from the error text. harness.py fails loudly instead.

Usage — Ollama

bash
ollama run hf.co/textilelabs/Loom-Spark-3.2 "who are you"

template and params are read automatically. Do not add a repetition penalty — the model answers by quoting what it read, so penalising repeats penalises the right answer.

Usage — transformers

python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tok = AutoTokenizer.from_pretrained("textilelabs/Loom-Spark-3.2")
model = AutoModelForCausalLM.from_pretrained("textilelabs/Loom-Spark-3.2").eval()
eot = tok.convert_tokens_to_ids("<|eot|>")

def ask(message, tools=False):
    p = f"<tools:{'on' if tools else 'off'}>\n<user>\n{message}\n<|eot|>\n<loom>\n"
    ids = tok(p, return_tensors="pt", add_special_tokens=False).input_ids
    with torch.no_grad():
        out = model.generate(ids, max_new_tokens=96, do_sample=False, eos_token_id=eot,
                             pad_token_id=tok.convert_tokens_to_ids("<|pad|>"))[0]
    return tok.decode(out[ids.shape[1]:], skip_special_tokens=False).replace("<|eot|>", "").strip()

Prompt format is exact: <tools:off>\n<user>\n{message}\n<|eot|>\n<loom>\n. For more turns, append {reply}<|eot|>\n<user>\n{next message}\n<|eot|>\n<loom>\n.

How it was built

architectureLlama — 20 layers × 320d, FFN 864, GQA (5 heads / 1 KV), SwiGLU, RoPE, tied embeddings
parameters22,827,840
context2,048
vocabulary4,096 custom BPE (the Spark-line tokenizer, unchanged from Spark 3)
optimiserMuon (0.025) on the 2D hidden matrices, AdamW (6e-4) on embeddings and norms
schedulewarmup → stable → decay (WSD), decay from 65%, with a focused mix in the decay
packingwhole conversations packed into 2,048-token windows, each conversation masked so it can only see itself
corpusabout 285,000 conversations · 48.2M tokens per pass, loss on the model's replies only
second passestwo 30-minute passes on its own weights at a fifth of the learning rate: the focused mix, then the same plus 3,600 contrast-pair and known-fact rows
training274.9M tokens · 12.0 tokens per parameter · from random init
hardwareKaggle, 2 × NVIDIA T4 · 2 h 13 min (74 min, then 30 + 30 min)

The per-conversation mask mattered more than anything else. Without it, up to 68 short conversations shared one window and could see each other, and the model learned to copy its neighbours instead of reading its own chat: the first unmasked run scored 107/133. Same data with the mask: 120/133.

Files

config.json / model.safetensors           the model
tokenizer.json / tokenizer_config.json    custom BPE tokenizer, 4,096 tokens
loom-spark-3.2-f16.gguf                   for Ollama / llama.cpp
harness.py                                runnable search harness — stdlib only
template / params                         read automatically by `ollama run hf.co/...`
Modelfile                                 for building locally
ATTRIBUTION.md                            required credits for the training corpora

Training data

slicesource
grounded reading, three-paragraph reading, and "the result doesn't say"SQuAD 2.0 (CC BY-SA 4.0)
multi-hop and trivia readingHotpotQA (CC BY-SA 4.0) · TriviaQA (Apache 2.0) · Wikipedia (CC BY-SA)
when to reach for a toolMASSIVE (CC BY 4.0) · CLINC150 (CC BY 3.0)
instruction followingdatabricks-dolly-15k (CC BY-SA 3.0)
multi-turn dialogue structureOpenAssistant OASST1 (Apache 2.0)
identity, limits, declines, warmth, memory within a chat, injection resistanceTextile Labs — written for Loom

Every search query is derived mechanically from these sources. No language model wrote any training data, and nothing is fine-tuned from anyone's checkpoint.

License

Model: MIT. Training data retains its original licences and attribution.