Matias/lora-kernel-distributor-wiki-gemma4-e4b
distributor-wiki@v2 — a LoRA that navigates a wiki instead of memorising it
Research preview from [lora-kernel](https://github.com/EvolvingAgentsLabs/lora-kernel). This adapter isn't a general assistant, and it holds no facts. It learned one habit: walking a library of short markdown pages to answer a question. It searches, opens a page, opens the one section it needs, follows the link written inside that sentence, sends arithmetic to a calculator, and cites the statement its answer rests on. When the library holds nothing on the question, it says Not in my library.
The weights hold the navigation and the markdown holds the content. When a fact changes, you edit a page and leave the model as it is.
Try it

The notebook asks six two-hop questions to the bare model and then to this adapter, then changes one number in one markdown page and asks again. On an L4 on 2026-09-26 [ran], the bare model scored 2/6 and the adapter 6/6. After a supplier's lead time was edited from 7 to 14 days, the adapter answered 14 days [rfu§lead-time] with no retraining. It needs a bf16 GPU and no login.
What it was measured on [ran]
The evaluation world is a generated distributor wiki (knowledge/distributor-wiki, seed 20260924). The adapter never saw that world in training: its corpus was drawn from 32 other generated worlds. An answer counts only if the value is right and the cited section actually says it.
Runs: results/M7-W9-atomic-statements-20260924, B1, and results/B5-comparison-corpus-20260926 in the repository. The release manifest is releases/distributor-wiki@v2.json, and the adapter sha256 is d4d91ea4aa7a8ffcd82beabb87d07112c02b212388412fb61f43df49d3eb93d6.
Limits, stated plainly
- Every world it has seen or been scored on is generated by one generator. It hasn't been measured on a hand-written or real organisation's library. A model trained on generated data can learn the generator's quirks, and this one hasn't been tested outside them.
- It expects this page format: YAML front matter (
id,when,what) followed by one§anchorstatement per line, with links as[[page-id]]inside the statements. Seeknowledge/distributor-wiki/. - It needs the runtime. The model writes
<search>,<open>and<calc>tags. The runtime stops at each closing tag, writes the result inline after=, and lets generation continue. This is not OpenAItool_calls: one adapter in this project scored 11/90 served that way and 90/90 served inline as it was trained. - Thinking is off (
enable_thinking=false) in every render. The vision and audio towers are excluded from the LoRA.
How to run it
The runtime lives in memory/ (Library, Conversation). The runner that produced the numbers above is training/wiki/wiki_arm.py, served through vLLM with LoRA enabled. To reproduce a score, see docs/SERVING.md and docs/MEMORY.md §1.6 in the repository.
Recipe: r 16, alpha 32, 3 epochs, lr 2e-4, effective batch 16, maxseq 1536, all attention and MLP projections, seed 1. The corpus is `training/wiki/data/traincmp.jsonl (728 walks, sha256 7b307d53…`).
Apache-2.0, like the base model.
