Team Ai
Modelpublic

cimo001/embeddinggemma-300m

sourceHugging Facemitupdated 27d agoView on Hugging Face
1likes52downloads
Model Card

ONNX - embeddinggemma-300m

FP32 = model.onnx, model.onnx_data

INT8 = modelquantized.onnx, modelquantized.onnx_data

Usage

# onnxruntime-gpu for run it on GPU
pip install onnxruntime sentencepiece numpy
  • —src/helper.py = onnx session builder (provider selection, threads, memory options)
  • —src/example.py = full pipeline: sentencepiece tokenization, batching, inference, cosine similarity
python3 src/example.py
0.600231 | The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear, is a bear species endemic to China.
0.153004 | hi
0.554038 | パンダはクマ科の哺乳類で、中国の固有種である。

Note:

  • —No transformers/tokenizers dependency: the tokenizer is tokenizer.model directly with the native ids<br> (bosid, eosid, pad_id from the sentencepiece model).
  • —Prompt prefix per mode:<br> document = title: none | text: {text}<br> query = task: search result | query: {text}
  • —Input format: [bos] text [eos], padded per batch to the longest sequence.
  • —Output = sentence_embedding, 768 dimensions.
  • —Normalize input with NFKC for consistent multilingual scores.