Team Ai
Modelpublic

OO-LD/oold-quantities-lean-phi4-r16-s1000

sourceHugging Faceapache-2.0updated 6d agoView on Hugging Face
0likes17downloads
Model Card

oold-lean-phi4-r16-s1000

A LoRA for schema-constrained quantity extraction, trained and measured by oold-llm-bench. Size series, another architecture.

It is worth having in one condition. Where the prompt carries the class catalogue, every adapter measured so far scores below the model it was trained from. Where the prompt carries nothing and a grammar does the constraining, this is what it buys:

conditionbasethis adapter
no-catalog-enforced0.3880.733 (n=119)
no-catalog-not-enforced0.0000.080 (n=120)
schema-prose-catalog-enforced0.8560.883 (n=120)

Recipe

basemicrosoft/phi-4
rank16
target_modulesall-linear
steps1000
examples1,000
corpusgenerated quantity documents, 100 classes, QUDT units
evaluated onWiki-Measurements, held-out class half, human prose

Trained on machine prose and scored on human prose, deliberately: the register is held out so a score is not a report on the generator.

Serving

vLLM, loaded beside the base so the pair differs in the adapter and nothing else:

vllm serve microsoft/phi-4 --enable-lora --lora-modules oold=OO-LD/oold-lean-phi4-r16-s1000 --max-lora-rank 16

--max-lora-rank has to be at least the rank, or the adapter loads and changes nothing. An adapter loaded at runtime through POST /v1/load_lora_adapter does not survive scale-to-zero.

What it does not do

The catalogue is worth more than the tune. Untuned with the class list in the prompt the same base reaches 0.873 on this task, above every adapter here, so this is a way to buy a short prompt and not a way to buy accuracy.

Across 4B to 32B in two families, parameter count orders neither the base score nor the gain. The numbers and the method are on the benchmark.