Team Ai
Modelpublic

ytu-ce-cosmos/modernbert-tr-embed

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
6likes902downloads
Model Card

<p align="center"> <img src="assets/logo.webp" width="20%" alt="ModernBERT-TR Embed" /> </p> <h1 align="center">ModernBERT-TR Embed</h1>

A 150M-parameter Turkish text-embedding model.

Results

ModelParamsRetrClassifPairClsClusterSTSBitext**Mean**
ModernBERT-TR-Embed (ours)150M59.476.569.263.377.694.168.14
ytu-ce-cosmos/turkish-e5-large560M61.572.662.860.980.099.367.17
microsoft/harrier-oss-v1-0.6b600M60.171.158.663.374.598.665.57
intfloat/multilingual-e5-large560M61.769.265.660.881.099.066.56
Qwen/Qwen3-Embedding-4B4B63.170.260.161.377.097.866.69

How was this model trained?

  1. 1.We embedded ~7.9M Turkish passages with the teacher, then trained our model to reproduce those embeddings. A projector maps the teacher's 4096-d vectors down to our model's 768 dimensions. Following Jasper/Stella distillation recipe, a three-term loss aligns the embeddings from both: a per-passage cosine loss, a similarity-matrix loss matching the student's and teacher's Gram matrices within batch, and a CoSENT-style hinge that reproduces the teacher's pairwise-similarity ordering.
  1. 1.From the distilled model we ran three fine-tunings, all supervised by teacher embeddings:
  2. 2.Retrieval: the student ranks the correct passage above hard negatives for a given query. Trained with an InfoNCE contrastive loss over in-batch and hard negatives, plus a KL term matching the teacher's softmax ranking over each query's candidates.
  3. 3.Multi-task: the retrieval objective plus Turkish language-understanding tasks: NLI, STS as in CoSENT on teacher cosine, supervised contrastive classification as in SupCon, and a replay of the cosine distillation on classification text.
  4. 4.Cross-lingual: the multi-task recipe with rebalanced task weights and added English retrieval passages, to improve English-Turkish alignment.
  1. 1.We weight-average the checkpoints described above into a single model.

Usage

sentence-transformers

python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("mrbesher/modernbert-tr-embed")

emb = model.encode(["bir cümle", "başka bir cümle"], normalize_embeddings=True)

q = model.encode(["soru"], prompt_name="query", normalize_embeddings=True)
d = model.encode(["döküman"], normalize_embeddings=True)

The retrieval (see config_sentence_transformers.json): Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery:{text}.

ONNX Runtime

The onnx/ folder has the grapgh for the token embeddings, it includes 3 graphs: token embeddings, mean-pool and L2-norm.

python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("mrbesher/modernbert-tr-embed", backend="onnx",
                            model_kwargs={"file_name": "onnx/model_fp16.onnx"})

Text Embeddings Inference (TEI)

TEI is L2-normalizes by default. Pass per-request prompts for retrieval queries.

bash
text-embeddings-router --model-id mrbesher/modernbert-tr-embed --dtype float16

Encoderfile

See the companion encoderfile repo.

Training data

Turkish retrieval (msmarco-tr, Squad-TR train/dev), Turkish NLI (boun-tabi/nli_tr train), Turkish STS-B (train), and Turkish classification-domain text (product reviews, news, social), all teacher-supervised. We check for leaks with text-hash against every MTEB(Turkish) test split.

License & attribution

  • —License: apache-2.0.