Menterium/NuExtract-1.5-tiny-onnx-cpu-int4
NuExtract-1.5-tiny — ONNX Runtime GenAI (CPU, int4)
ONNX export of numind/NuExtract-1.5-tiny for ONNX Runtime GenAI, quantized to int4 for CPU inference. Menterium converted it for in-process use in a .NET desktop application. Apart from the int4 quantization the weights are unchanged. NuMind does not publish an ONNX export of this model.
This is an unofficial conversion and is not affiliated with or endorsed by NuMind.
Model
Usage
NuExtract is not a chat model. It receives an empty JSON template and a text, and fills the template with values copied from the text ("" when a value is not found):
<|input|>
### Template:
{
"Name": "",
"City": ""
}
### Text:
Max Mustermann, Musterstr. 15-17, 12345 Musterstadt
<|output|><|input|>, <|output|> and the closing <|end-output|> are plain text in this tokenizer, exactly as in the original model; they are not special tokens. A JSON schema (constrained decoding) keeps the output valid and ends the generation after the closing brace.
Python:
import json
import onnxruntime_genai as og
model = og.Model("NuExtract-1.5-tiny-onnx-cpu-int4")
tokenizer = og.Tokenizer(model)
prompt = (
"<|input|>\n### Template:\n{\n \"Name\": \"\",\n \"City\": \"\"\n}\n"
"### Text:\nMax Mustermann, Musterstr. 15-17, 12345 Musterstadt\n\n<|output|>"
)
tokens = tokenizer.encode(prompt)
# Constrained decoding: the output must match this JSON schema.
schema = json.dumps({
"type": "object",
"properties": {"Name": {"type": "string"}, "City": {"type": "string"}},
"required": ["Name", "City"],
"additionalProperties": False,
})
params = og.GeneratorParams(model)
params.set_search_options(max_length=len(tokens) + 256)
params.set_guidance("json_schema", schema)
generator = og.Generator(model, params)
generator.append_tokens(tokens)
while not generator.is_done():
generator.generate_next_token()
print(tokenizer.decode(generator.get_sequence(0)[len(tokens):]))
# {"Name": "Max Mustermann", "City": "Musterstadt"} (printed indented, one field per line)Conversion
Model Builder of onnxruntime-genai 0.16.0 with transformers 5.17.0 and PyTorch 2.14.1 (CPU) on Python 3.12. The 0.16 model builder needs transformers 5.x; with 4.x it fails on import. requirements-builder.txt in this repository is the complete frozen environment (pip freeze, 49 packages including all transitive dependencies) that reproduces this export; install from it into a fresh virtual environment.
The source revision is pinned: download exactly that commit and let the builder read it from disk (-i); with -m alone the builder fetches whatever is current on main.
pip install -r requirements-builder.txt --extra-index-url https://download.pytorch.org/whl/cpu
hf download numind/NuExtract-1.5-tiny --revision 63e2e80c804d9c97f3f19a4aa25613e7beca83c9 --local-dir NuExtract-1.5-tiny-src
python -m onnxruntime_genai.models.builder -m numind/NuExtract-1.5-tiny -i NuExtract-1.5-tiny-src -o NuExtract-1.5-tiny-onnx-cpu-int4 -p int4 -e cpu --extra_options hf_token=falseRebuilding this way reproduces model.onnx, model.onnx.data, genai_config.json, tokenizer.json and chat_template.jinja byte for byte (checksums above). tokenizer_config.json differs only in the metadata flag is_local (true when the source is read from disk), which does not change tokenization.
Notes
- Quantization to int4 costs some accuracy compared with the original model. On German order confirmations we also tested an fp32 export: it was not more accurate in practice but about four times slower.
- Speed on an Intel Core i7-9700K (8 cores, AVX2, no GPU): roughly 240 tokens/s for the prompt and 25 tokens/s for the answer.
- Small models benefit from short templates. We combine this model with deterministic rules and only ask it for the fields the rules do not find.
Files (SHA-256)
License
MIT, like the original model (see `LICENSE`). NuExtract-1.5-tiny is based on Qwen2.5-0.5B, which is licensed under the Apache License 2.0 (see `LICENSE-Qwen2.5`).
