Team Ai
Modelpublic

TREEIndustries/Llama-3-ARDY-Text-Encoder-ONNX

sourceHugging Facellama3updated 2mo agoView on Hugging Face
1likes
Model Card

Llama 3 ARDY Text Encoder ONNX

Built with Meta Llama 3.

This is the reusable INT4 ONNX text encoder for NVIDIA ARDY runtime motion prompting. It converts a raw prompt token sequence into the 4096-dimensional text feature consumed by every Core and G1 ARDY denoiser in the TREE Industries ONNX releases:

The graph combines:

  1. 1.meta-llama/Meta-Llama-3-8B-Instruct
  2. 2.the McGill NLP LLM2Vec MNTP adapter
  3. 3.the McGill NLP LLM2Vec supervised adapter
  4. 4.bidirectional Llama attention and ARDY-compatible mean pooling

The two adapters were merged before ONNX export. Constant-weight matrix multiplications were quantized to symmetric INT4 MatMulNBits operators with block size 32. The graph retains FP16 activations and returns an FP32 pooled embedding.

Files

  • —onnx/text_encoder_int4.onnx is the ONNX graph.
  • —onnx/text_encoder_int4.onnx.data contains its external INT4 weights.
  • —tokenizer/tokenizer.json and tokenizer/tokenizer_config.json are the standard Llama 3 tokenizer assets.
  • —tokenizer/tokenizer.ardy is a compact case-sensitive tokenizer asset for native runtimes.
  • —validation/text_encoder_reference.json contains five FP16 reference embeddings and their exact tokenized inputs.
  • —SHA256SUMS.txt provides integrity hashes.

ONNX interface

Inputs are fixed to a 64-token sequence:

NameTypeShapeMeaning
input_idsint64[1, 64]Left-padded Llama 3 token IDs
attention_maskint64[1, 64]Token attention mask
embed_maskint64[1, 64]Tokens included in mean pooling

Output:

NameTypeShape
text_embeddingfloat32[1, 4096]

Prompts are formatted using the Llama 3 user header and end-of-turn token before tokenization. Padding is on the left; truncation is on the right.

Validation

The graph passed full ONNX checking and was executed with ONNX Runtime 1.20.1 using DmlExecutionProvider on an NVIDIA GeForce RTX 4080 Laptop GPU.

Across five stored FP16 reference prompts:

  • —minimum cosine similarity: 0.9772086
  • —mean cosine similarity: 0.9805653
  • —maximum cosine similarity: 0.9823330
  • —every output had shape [1, 4096]
  • —every output value was finite

Detailed per-prompt results are in validation/VALIDATION.md.

Source provenance

ComponentSource revision
meta-llama/Meta-Llama-3-8B-Instruct8afb486c1db24fe5011ec46dfbe5b5dccdb575c2
McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp31474e395ada192e8ed1586db6be79fb3b70c9c0
McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-supervisedbaa8ebf04a1c2500e61288e7dad65e8ae42601a7

Licenses and required notices

The Llama Materials and this derivative model are distributed under the Meta Llama 3 Community License. Use must also comply with the Meta Llama 3 Acceptable Use Policy.

Required Meta attribution:

Meta Llama 3 is licensed under the Meta Llama 3 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.

The LLM2Vec adapters and project code are MIT licensed. See LICENSES/MIT-LLM2Vec.txt and NOTICE.

LLM2Vec citation

bibtex
@inproceedings{llm2vec,
  title={LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders},
  author={BehnamGhader, Parishad and Adlakha, Vaibhav and Mosbach, Marius and Bahdanau, Dzmitry and Chapados, Nicolas and Reddy, Siva},
  booktitle={First Conference on Language Modeling},
  year={2024},
  url={https://openreview.net/forum?id=IW1PR7vEBf}
}