OO-LD/oold-quantities-lean-phi4-r16-s1000
oold-lean-phi4-r16-s1000
A LoRA for schema-constrained quantity extraction, trained and measured by oold-llm-bench. Size series, another architecture.
It is worth having in one condition. Where the prompt carries the class catalogue, every adapter measured so far scores below the model it was trained from. Where the prompt carries nothing and a grammar does the constraining, this is what it buys:
Recipe
Trained on machine prose and scored on human prose, deliberately: the register is held out so a score is not a report on the generator.
Serving
vLLM, loaded beside the base so the pair differs in the adapter and nothing else:
vllm serve microsoft/phi-4 --enable-lora --lora-modules oold=OO-LD/oold-lean-phi4-r16-s1000 --max-lora-rank 16--max-lora-rank has to be at least the rank, or the adapter loads and changes nothing. An adapter loaded at runtime through POST /v1/load_lora_adapter does not survive scale-to-zero.
What it does not do
The catalogue is worth more than the tune. Untuned with the class list in the prompt the same base reaches 0.873 on this task, above every adapter here, so this is a way to buy a short prompt and not a way to buy accuracy.
Across 4B to 32B in two families, parameter count orders neither the base score nor the gain. The numbers and the method are on the benchmark.
