ConfidentialMind/SampoTron
SampoTron — cross-lingual query adapter for Nemotron-3-Embed-1B
SampoTron is a cross-lingual query adapter for the base model nvidia/Nemotron-3-Embed-1B-BF16. It improves cross-lingual retrieval between English, Finnish, and Swedish while leaving same-language retrieval untouched through language routing.
This repository ships two deployment forms:
model.safetensorsat the root: the standalone merged query encoder (bfloat16), loadable directly as a SentenceTransformer.merge_manifest.jsonrecords the merge (base revision, adapter sha256, dtype).adapter/: the rank-64 query-only LoRA (84 MB) for PEFT users; applying it to the pinned base revision reproduces the merged model up to bfloat16 rounding.
When to use it
Two deployment paths — pick one, never stack them. The merged model.safetensors at the root already has the LoRA folded in. adapter/ applies the same LoRA to a pristine base. Either route cross-language queries to the merged model, or load the base model and attach adapter/; never attach adapter/ to the merged model, or the LoRA effect is applied twice.
Use the adapter only for cross-language queries — the query language differs from the index language. Encode same-language queries with the unmodified base model. Documents are always encoded by the unmodified base model, so existing base-model document indexes work unchanged.
Routing requires the application to know the query and index languages; use sampotron when they differ and nemotron-base otherwise. Serving geometry is unchanged: encode_query, Matryoshka prefix slice from the native 2048 dimensions to the first 1024, L2 renormalize after slicing (served 1024-d).
vLLM serving
The adapter can be served alongside the pristine base in a single vLLM process, with per-request routing through the model field. vLLM ≥ 0.29, --max-lora-rank 64 required.
# 1. fetch the LoRA adapter (84 MB; model.safetensors stays on the Hub)
hf download ConfidentialMind/SampoTron --include 'adapter/*' --local-dir ./sampotron-adapter
# 2. serve base + adapter in one process
vllm serve nvidia/Nemotron-3-Embed-1B-BF16 \
--runner pooling \
--served-model-name nemotron-base \
--host 127.0.0.1 --port 8011 \
--max-model-len 1024 \
--enable-lora --max-lora-rank 64 \
--lora-modules sampotron=./sampotron-adapter/adapter \
--hf-overrides '{"is_matryoshka": true, "matryoshka_dimensions": [1024, 2048]}' \
--pooler-config '{"dimensions": 1024}'The --hf-overrides/--pooler-config pair makes the server apply the Matryoshka prefix slice to 1024 dimensions and L2-normalize, so responses are the served geometry directly. Route by request:
Cross-language query (adapter route):
{"model": "sampotron", "input": ["query: how to dispute an invoice charge"]}Same-language query, and every document (base route):
{"model": "nemotron-base", "input": ["query: yrityksen arvonlisävero"]}{"model": "nemotron-base", "input": ["passage: <document text>"]}Notes:
- The command serves the base repository's
main. The snapshot this adapter was merged against is recorded inmerge_manifest.json; pass--revisionwith that value to pin it. - The
--pooler-configslice replaces the client-side 2048→1024 slice; use one or the other, never both. - When serving the merged model directly, use mean pooling; last-token pooling returns incorrect embeddings.
Results (nDCG@10)
Routing keeps every same-language and native-Finnish score exactly at the base model's values. All six cross-language directions improve, from +0.012 (en→sv) to +0.097 (fi→en).
The merged model measured 0.2900 versus 0.2908 for the live-adapter path, with max per-cell drift 0.0022, consistent with bfloat16 merge rounding; merge_manifest.json records the validation.
Caveats
- Use the merged model or LoRA adapter only for cross-language queries; use the base model for same-language queries.
License
The merged model and the LoRA adapter are redistributed under the OpenMDW License Agreement, version 1.1 (OpenMDW-1.1), the same terms as the base model. See LICENSE for the license text and NOTICE for the upstream origin notices.
