OO-LD/oold-quantities-lean
oold-quantities-lean What every adapter under this organisation was trained on. Chat-formatted examples of quantity extraction, generated from the quantity schemas and rendered to prose. The prompt is the evaluation prompt The system message is built by the same function that builds it at evaluation time, under the same arm, and is character-identical to it. Trained on one wording and evaluated on another, a result is a wording difference. Lean means the class… See the full description on the dataset page: https://huggingface.co/datasets/OO-LD/oold-quantities-lean.
oold-quantities-lean
What every adapter under this organisation was trained on. Chat-formatted examples of quantity extraction, generated from the quantity schemas and rendered to prose.
The prompt is the evaluation prompt
The system message is built by the same function that builds it at evaluation time, under the same arm, and is character-identical to it. Trained on one wording and evaluated on another, a result is a wording difference.
Lean means the class catalogue is shown and the schema is not. Trained with a full schema in context, a model learns to copy from the prompt, and the tuned-with-no-schema condition the adapters exist for would measure nothing.
Held out by class, and by register
Roughly half the describable classes are excluded from training entirely. At evaluation the offered catalogue still carries both halves, or a held-out class could not be chosen and the contrast could not be measured.
The register is held out too. These documents are generated prose; the adapters are scored on Wikipedia sentences. A number measured this way is not a report on the generator.
Seeds are disjoint from every evaluation draw, so no scored document was trained on.
Every example round-trips
Each assistant turn is fed back through the real extractor and the real grader and must score 1.00. A training set that does not round-trip teaches a model answers the benchmark would mark wrong. 17 of 1,017 draws were rejected by that gate.
Files
Both carry a UTF-8 byte order mark. Kept rather than stripped, because these are the bytes the adapters were trained on; read them with utf-8-sig.
Regenerate with oold-llm-bench.
