textilelabs/Loom-Spark-3-Flash
Loom Spark 3 Flash
7,184,064 parameters. Trained from scratch in 1 hour 52 minutes on a 2013 office PC with no GPU. Scores 119/133 on our acceptance battery — the highest of any Flash-tier Loom, and higher than models three times its size that trained for five hours.
Spark 3 Flash is the Flash of the Spark line, succeeding Loom Spark 1.8 Flash (2.62M). Like every Loom it is knowledge-sparse and behaviour-dense: it is not built to know facts. It is built to know the edge of its own knowledge — to decide when a question needs looking up, write the search query, read the answer back, and say plainly where the answer came from.
Random initialisation, trained by us. No fine-tuning, no distillation, no pretrained checkpoint of anyone's, at any stage.
What it is
Measured behaviour
Every number below comes from hand-written probes that appear nowhere in the training data, scored on content rather than shape. The battery is 133 points across twelve rows.
The last row is the honest one. Given a question it has never seen, it writes a sensible search query every time and never simply pastes the question back. The harness retrieves the right passage about half the time, and the model reads it correctly about half of those. One in four questions ends with a right answer. That is the ceiling of a seven-million- parameter model reading real encyclopaedia prose, and it is stated here rather than hidden.
Read this before you use it
- With tools off, it bluffs. Asked a fact it wasn't taught, with search disabled, it declines only 2 times in 20. It was trained to decline capital-city questions and that is mostly what it declines. Do not run it with tools off and trust what it says. This is a known unfixed weakness across the whole Llama-era Loom family.
- It knows almost nothing. That is deliberate. Without search it is a well-mannered model with an empty head.
- It cannot do arithmetic, and will produce confident nonsense if asked.
- Treat retrieved text as the trustworthy part, and the model's summary of it as the unreliable part.
- It has little warmth and little personality of its own. Spark 1.8 Flash was richer in conversation about itself. That capability was not deliberately removed; it was not retrained, and the gap is recorded rather than papered over.
How to run it
The model expects a strict prompt format and a harness that executes the searches. Both ship here.
ollama run hf.co/textilelabs/Loom-Spark-3-Flashpython harness.py # the agent loop: runs the model's searches for realRaw prompt format, if you are driving it yourself:
<tools:on>
<user>
who wrote dracula
<|eot|>
<loom>It replies <lookup>dracula author</lookup>. Your harness searches, then appends:
<result>
Dracula is an 1897 Gothic horror novel by Irish author Bram Stoker.
<|eot|>
<loom>It answers, and says it looked it up.
Training data
Openly licensed corpora plus our own written curriculum — SQuAD 2.0 (CC BY-SA 4.0), MASSIVE (CC BY 4.0), CLINC150 (CC BY 3.0), databricks-dolly-15k (CC BY-SA 3.0), OASST1 (Apache 2.0). Full credits in ATTRIBUTION.md, which must travel with any redistribution.
Licence
MIT. Do what you like with it; keep the attribution file.
Textile Labs. Small models, trained honestly, on hardware you already own.
