Horizon-Labs/multilingual-reranker-small
Multilingual Reranker (small, 141M)
A cross-encoder reranker for search and RAG in 40+ languages, built on mmBERT-small. Give it a query and candidate passages, for example the top 20-100 hits of a vector or keyword search. It scores each pair, so you can re-sort the candidates and keep the best ones for your LLM. Queries and passages can be in different languages. It is distilled from Qwen3-Reranker-4B into a model about 28x smaller. Apache-2.0, trained on openly licensed web text. ONNX files for CPU and the browser (transformers.js) are included. A larger, more accurate version is available as Horizon-Labs/multilingual-reranker-base. Try it in the browser.
- One output logit per (query, passage) pair: higher = more relevant.
sigmoid(logit)gives a 0-1 relevance score. - Max length: trained with 384 tokens per pair. Longer passages are truncated; split long documents into chunks.
onnx/model_quantized.onnx(int8 embeddings, 268 MB): scores differ from fp32 by 0.001 on average (sigmoid scale) over 1240 benchmark pairs.
Usage
sentence-transformers:
from sentence_transformers import CrossEncoder
model = CrossEncoder("Horizon-Labs/multilingual-reranker-small")
query = "How tall is the Eiffel Tower?"
passages = ["The Eiffel Tower is 330 metres tall.", "La tour Eiffel a été construite pour l'Exposition universelle de 1889.",
"The Statue of Liberty is 93 metres tall."]
print(model.rank(query, passages)) # [{'corpus_id': 0, 'score': ...}, ...]transformers:
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tok = AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-small")
model = AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-small").eval()
enc = tok([query] * len(passages), passages, padding=True, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
scores = model(**enc).logits[:, 0].sigmoid()transformers.js:
import { AutoTokenizer, AutoModelForSequenceClassification } from "@huggingface/transformers";
const tok = await AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-small");
const model = await AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-small", { dtype: "q8" });
const enc = await tok(new Array(passages.length).fill(query), { text_pair: passages, padding: true, truncation: true });
const { logits } = await model(enc);Evaluation
nDCG@10 for reranking the candidate lists of public reranking benchmarks (the MTEB versions). The benchmarks were used only for evaluation, never for training or model selection. Every model sees the same queries and candidates, through the same script (code/). MIRACL: 60 queries per language with 100 candidates each. Wikipedia (WikipediaRerankingMultilingual): 60 queries per language with 9 candidates each. "Other" = ESCI (es, jp, us), RuBQ, T2Reranking, mMARCO-ja and AskUbuntu.
- The table shows the released checkpoint. Means over two training seeds: small MIRACL .730 / Wikipedia .959 / other .783 / mean .824; base .760 / .964 / .796 / .840.
- For reference, ranking the same candidates by bge-m3 dense-embedding similarity alone scores 0.797 / 0.927 / 0.788 (mean 0.837), so on this test a reranker adds most on top of a weaker first-stage retriever (BM25 or a small embedding model). Full comparison: leaderboard.
- Models ahead of this one: MIRACL: bge-reranker-v2-m3, Qwen3-Reranker-0.6B, Qwen3-Reranker-4B; Wikipedia: Qwen3-Reranker-4B; other: bge-reranker-v2-m3, Qwen3-Reranker-0.6B, Qwen3-Reranker-4B. bge-reranker-v2-m3 was trained on MIRACL's training set; we were not.
Per set (nDCG@10):
Training
- Passages: about 400k passages of 2-6 sentences from FineWeb-2 and FineWeb (ODC-BY) in 47 languages.
- Queries: Qwen3.8-27B (Apache-2.0) wrote a natural question and a keyword query for each passage, in the passage's language, plus English questions for 15% of the non-English passages (cross-lingual search).
- Scale: 334,400 queries with 16 candidates each (5.35M teacher-scored pairs), sampled evenly across languages. The data is published as Horizon-Labs/multilingual-rerank-distill.
- Schedule: 1 epoch, learning rate 5e-5, 16 queries x 16 candidates per step, max 384 tokens per pair.
- Candidates: for each query, the 15 most similar passages in the same language by bge-m3 dense retrieval (hard negatives) plus the source passage.
- Labels: Qwen3-Reranker-4B (Apache-2.0) scored every (query, candidate) pair. The student learns the teacher's ranking with a listwise KL loss over each query's 16 candidates, plus a pointwise loss on the teacher's relevance probability. The teacher's scores are soft labels, so the near-duplicate passages that dense retrieval finds are not wrongly treated as negatives.
- Code:
code/in this repository.
Limitations
- The model inherits the teacher's judgement, including its mistakes. It is weaker than the teacher, especially on long or technical passages.
- Training queries are LLM-written questions and keyword queries over web text. Very domain-specific search (legal, medical, code) and conversational queries are less covered.
- Pairs are truncated at the max length. Rerank passage-sized chunks, not whole documents.
