Team Ai
Modelpublic

Menterium/NuExtract-1.5-tiny-onnx-cpu-int4

sourceHugging Facemitupdated 8h agoView on Hugging Face
0likes
Model Card

NuExtract-1.5-tiny — ONNX Runtime GenAI (CPU, int4)

ONNX export of numind/NuExtract-1.5-tiny for ONNX Runtime GenAI, quantized to int4 for CPU inference. Menterium converted it for in-process use in a .NET desktop application. Apart from the int4 quantization the weights are unchanged. NuMind does not publish an ONNX export of this model.

This is an unofficial conversion and is not affiliated with or endorsed by NuMind.

Model

Original modelnumind/NuExtract-1.5-tiny, revision 63e2e80c804d9c97f3f19a4aa25613e7beca83c9, MIT
Base model of the originalQwen/Qwen2.5-0.5B, Apache 2.0
Parametersabout 0.5 billion
FormatONNX, int4 weight-only quantization (MatMulNBits), for the CPU execution provider
Sizeabout 0.33 GB
Context length32,768 tokens (genai_config.json)
Runtimeonnxruntime-genai 0.16 or newer (NuGet Microsoft.ML.OnnxRuntimeGenAI, PyPI onnxruntime-genai)

Usage

NuExtract is not a chat model. It receives an empty JSON template and a text, and fills the template with values copied from the text ("" when a value is not found):

<|input|>
### Template:
{
    "Name": "",
    "City": ""
}
### Text:
Max Mustermann, Musterstr. 15-17, 12345 Musterstadt

<|output|>

<|input|>, <|output|> and the closing <|end-output|> are plain text in this tokenizer, exactly as in the original model; they are not special tokens. A JSON schema (constrained decoding) keeps the output valid and ends the generation after the closing brace.

Python:

python
import json

import onnxruntime_genai as og

model = og.Model("NuExtract-1.5-tiny-onnx-cpu-int4")
tokenizer = og.Tokenizer(model)

prompt = (
    "<|input|>\n### Template:\n{\n    \"Name\": \"\",\n    \"City\": \"\"\n}\n"
    "### Text:\nMax Mustermann, Musterstr. 15-17, 12345 Musterstadt\n\n<|output|>"
)
tokens = tokenizer.encode(prompt)

# Constrained decoding: the output must match this JSON schema.
schema = json.dumps({
    "type": "object",
    "properties": {"Name": {"type": "string"}, "City": {"type": "string"}},
    "required": ["Name", "City"],
    "additionalProperties": False,
})

params = og.GeneratorParams(model)
params.set_search_options(max_length=len(tokens) + 256)
params.set_guidance("json_schema", schema)
generator = og.Generator(model, params)
generator.append_tokens(tokens)
while not generator.is_done():
    generator.generate_next_token()

print(tokenizer.decode(generator.get_sequence(0)[len(tokens):]))
# {"Name": "Max Mustermann", "City": "Musterstadt"}  (printed indented, one field per line)

Conversion

Model Builder of onnxruntime-genai 0.16.0 with transformers 5.17.0 and PyTorch 2.14.1 (CPU) on Python 3.12. The 0.16 model builder needs transformers 5.x; with 4.x it fails on import. requirements-builder.txt in this repository is the complete frozen environment (pip freeze, 49 packages including all transitive dependencies) that reproduces this export; install from it into a fresh virtual environment.

The source revision is pinned: download exactly that commit and let the builder read it from disk (-i); with -m alone the builder fetches whatever is current on main.

pip install -r requirements-builder.txt --extra-index-url https://download.pytorch.org/whl/cpu
hf download numind/NuExtract-1.5-tiny --revision 63e2e80c804d9c97f3f19a4aa25613e7beca83c9 --local-dir NuExtract-1.5-tiny-src
python -m onnxruntime_genai.models.builder -m numind/NuExtract-1.5-tiny -i NuExtract-1.5-tiny-src -o NuExtract-1.5-tiny-onnx-cpu-int4 -p int4 -e cpu --extra_options hf_token=false

Rebuilding this way reproduces model.onnx, model.onnx.data, genai_config.json, tokenizer.json and chat_template.jinja byte for byte (checksums above). tokenizer_config.json differs only in the metadata flag is_local (true when the source is read from disk), which does not change tokenization.

Notes

  • —Quantization to int4 costs some accuracy compared with the original model. On German order confirmations we also tested an fp32 export: it was not more accurate in practice but about four times slower.
  • —Speed on an Intel Core i7-9700K (8 cores, AVX2, no GPU): roughly 240 tokens/s for the prompt and 25 tokens/s for the answer.
  • —Small models benefit from short templates. We combine this model with deterministic rules and only ask it for the fields the rules do not find.

Files (SHA-256)

FileSize (bytes)SHA-256
genai_config.json1,5889ac771ba54d19d5037be767535fd7a08507eb1609b19d1a2aafef3075fdf43ac
model.onnx191,7254bab94c84fbdd6f3ef6351505aa3cad6d8011e99c5886b2e3def4ddd99f74cd3
model.onnx.data320,970,7524ed9defd489bb5e3454f643d46308c8ba27d3caf0f0e9982b0222f500c46acc4
tokenizer.json11,421,991be551d7071fd08ce0f7a7d30597b74565b1a62933188363ae7509af6529b2338
tokenizer_config.json866539e82d78f58ec889c58f779e8cb717d675bd8fadddd0260d58dbce245f8da0d
chat_template.jinja2,4817215110d1460d6dfdb8ed5bbc3f18b1772cc270404795c331bc0adc266a75008

License

MIT, like the original model (see `LICENSE`). NuExtract-1.5-tiny is based on Qwen2.5-0.5B, which is licensed under the Apache License 2.0 (see `LICENSE-Qwen2.5`).