Team Ai
Modelpublic

Horizon-Labs/multilingual-reranker-small

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes252downloads
Model Card

Multilingual Reranker (small, 141M)

A cross-encoder reranker for search and RAG in 40+ languages, built on mmBERT-small. Give it a query and candidate passages, for example the top 20-100 hits of a vector or keyword search. It scores each pair, so you can re-sort the candidates and keep the best ones for your LLM. Queries and passages can be in different languages. It is distilled from Qwen3-Reranker-4B into a model about 28x smaller. Apache-2.0, trained on openly licensed web text. ONNX files for CPU and the browser (transformers.js) are included. A larger, more accurate version is available as Horizon-Labs/multilingual-reranker-base. Try it in the browser.

  • —One output logit per (query, passage) pair: higher = more relevant. sigmoid(logit) gives a 0-1 relevance score.
  • —Max length: trained with 384 tokens per pair. Longer passages are truncated; split long documents into chunks.
  • —onnx/model_quantized.onnx (int8 embeddings, 268 MB): scores differ from fp32 by 0.001 on average (sigmoid scale) over 1240 benchmark pairs.

Usage

sentence-transformers:

python
from sentence_transformers import CrossEncoder

model = CrossEncoder("Horizon-Labs/multilingual-reranker-small")
query = "How tall is the Eiffel Tower?"
passages = ["The Eiffel Tower is 330 metres tall.", "La tour Eiffel a été construite pour l'Exposition universelle de 1889.",
            "The Statue of Liberty is 93 metres tall."]
print(model.rank(query, passages))   # [{'corpus_id': 0, 'score': ...}, ...]

transformers:

python
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tok = AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-small")
model = AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-small").eval()
enc = tok([query] * len(passages), passages, padding=True, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    scores = model(**enc).logits[:, 0].sigmoid()

transformers.js:

js
import { AutoTokenizer, AutoModelForSequenceClassification } from "@huggingface/transformers";
const tok = await AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-small");
const model = await AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-small", { dtype: "q8" });
const enc = await tok(new Array(passages.length).fill(query), { text_pair: passages, padding: true, truncation: true });
const { logits } = await model(enc);

Evaluation

nDCG@10 for reranking the candidate lists of public reranking benchmarks (the MTEB versions). The benchmarks were used only for evaluation, never for training or model selection. Every model sees the same queries and candidates, through the same script (code/). MIRACL: 60 queries per language with 100 candidates each. Wikipedia (WikipediaRerankingMultilingual): 60 queries per language with 9 candidates each. "Other" = ESCI (es, jp, us), RuBQ, T2Reranking, mMARCO-ja and AskUbuntu.

modellicenceMIRACL (18 languages)Wikipedia (16 languages)other (6 sets)meannote
this model (141M)Apache-2.00.7240.9590.7850.823multilingual
Horizon-Labs/multilingual-reranker-base (308M)Apache-2.00.7610.9640.7970.841multilingual
cross-encoder/ms-marco-MiniLM-L6-v2 (22M)Apache-2.00.3220.8420.6480.604English
BAAI/bge-reranker-base (278M)MIT0.6830.8580.7350.758Chinese/English
Alibaba-NLP/gte-reranker-modernbert-base (149M)Apache-2.00.4730.8980.7550.709English
BAAI/bge-reranker-v2-m3 (568M)Apache-2.00.8250.9580.8140.866multilingual; trained on MIRACL train
Qwen/Qwen3-Reranker-0.6B (596M)Apache-2.00.7850.9580.7930.845multilingual LLM reranker
Qwen/Qwen3-Reranker-4B (4B)Apache-2.00.8330.9710.8220.875multilingual LLM reranker; our teacher
  • —The table shows the released checkpoint. Means over two training seeds: small MIRACL .730 / Wikipedia .959 / other .783 / mean .824; base .760 / .964 / .796 / .840.
  • —For reference, ranking the same candidates by bge-m3 dense-embedding similarity alone scores 0.797 / 0.927 / 0.788 (mean 0.837), so on this test a reranker adds most on top of a weaker first-stage retriever (BM25 or a small embedding model). Full comparison: leaderboard.
  • —Models ahead of this one: MIRACL: bge-reranker-v2-m3, Qwen3-Reranker-0.6B, Qwen3-Reranker-4B; Wikipedia: Qwen3-Reranker-4B; other: bge-reranker-v2-m3, Qwen3-Reranker-0.6B, Qwen3-Reranker-4B. bge-reranker-v2-m3 was trained on MIRACL's training set; we were not.

Per set (nDCG@10):

setthis modelbge-reranker-v2-m3Qwen3-Reranker-4B (teacher)
AskUbuntu duplicate questions (English)0.6260.6790.700
ESCI product search, Spanish0.8460.8460.872
ESCI product search, Japanese0.8630.8610.869
ESCI product search, English0.8820.8950.908
MIRACL ar0.7580.8450.876
MIRACL bn0.7260.8930.896
MIRACL de0.6340.7480.791
MIRACL en0.6680.7280.790
MIRACL es0.6910.7810.767
MIRACL fa0.6870.8030.821
MIRACL fi0.7200.8380.856
MIRACL fr0.6420.7510.760
MIRACL hi0.6620.7430.781
MIRACL id0.5940.7220.641
MIRACL ja0.7640.8400.873
MIRACL ko0.7890.8410.856
MIRACL ru0.7270.8380.868
MIRACL sw0.8130.9040.859
MIRACL te0.7560.9610.912
MIRACL th0.7230.8780.887
MIRACL yo0.8830.9350.903
MIRACL zh0.7980.8090.851
mMARCO (Japanese)0.7340.8110.776
RuBQ (Russian)0.7970.8610.874
T2Reranking (Chinese)0.7480.7460.756
Wikipedia bg0.9490.9560.967
Wikipedia bn0.9590.9640.971
Wikipedia cs0.9940.9830.992
Wikipedia da0.9670.9550.988
Wikipedia de0.9590.9630.988
Wikipedia en0.9750.9830.986
Wikipedia fa0.9330.9530.956
Wikipedia fi0.9670.9730.988
Wikipedia hi0.9300.9170.914
Wikipedia it0.9550.9690.966
Wikipedia nl0.9530.9610.968
Wikipedia no0.9210.9310.945
Wikipedia pt0.9570.9490.978
Wikipedia ro0.9650.9420.982
Wikipedia sr1.0000.9730.983
Wikipedia sv0.9580.9620.969

Training

  • —Passages: about 400k passages of 2-6 sentences from FineWeb-2 and FineWeb (ODC-BY) in 47 languages.
  • —Queries: Qwen3.8-27B (Apache-2.0) wrote a natural question and a keyword query for each passage, in the passage's language, plus English questions for 15% of the non-English passages (cross-lingual search).
  • —Scale: 334,400 queries with 16 candidates each (5.35M teacher-scored pairs), sampled evenly across languages. The data is published as Horizon-Labs/multilingual-rerank-distill.
  • —Schedule: 1 epoch, learning rate 5e-5, 16 queries x 16 candidates per step, max 384 tokens per pair.
  • —Candidates: for each query, the 15 most similar passages in the same language by bge-m3 dense retrieval (hard negatives) plus the source passage.
  • —Labels: Qwen3-Reranker-4B (Apache-2.0) scored every (query, candidate) pair. The student learns the teacher's ranking with a listwise KL loss over each query's 16 candidates, plus a pointwise loss on the teacher's relevance probability. The teacher's scores are soft labels, so the near-duplicate passages that dense retrieval finds are not wrongly treated as negatives.
  • —Code: code/ in this repository.

Limitations

  • —The model inherits the teacher's judgement, including its mistakes. It is weaker than the teacher, especially on long or technical passages.
  • —Training queries are LLM-written questions and keyword queries over web text. Very domain-specific search (legal, medical, code) and conversational queries are less covered.
  • —Pairs are truncated at the max length. Rerank passage-sized chunks, not whole documents.