ytu-ce-cosmos/modernbert-tr-embed
<p align="center"> <img src="assets/logo.webp" width="20%" alt="ModernBERT-TR Embed" /> </p> <h1 align="center">ModernBERT-TR Embed</h1>
A 150M-parameter Turkish text-embedding model.
- Base model: `ytu-ce-cosmos/modernbert-tr-base`.
- Distilled from
Qwen/Qwen3-Embedding-8B.
Results
How was this model trained?
- We embedded ~7.9M Turkish passages with the teacher, then trained our model to reproduce those embeddings. A projector maps the teacher's 4096-d vectors down to our model's 768 dimensions. Following Jasper/Stella distillation recipe, a three-term loss aligns the embeddings from both: a per-passage cosine loss, a similarity-matrix loss matching the student's and teacher's Gram matrices within batch, and a CoSENT-style hinge that reproduces the teacher's pairwise-similarity ordering.
- From the distilled model we ran three fine-tunings, all supervised by teacher embeddings:
- Retrieval: the student ranks the correct passage above hard negatives for a given query. Trained with an InfoNCE contrastive loss over in-batch and hard negatives, plus a KL term matching the teacher's softmax ranking over each query's candidates.
- Multi-task: the retrieval objective plus Turkish language-understanding tasks: NLI, STS as in CoSENT on teacher cosine, supervised contrastive classification as in SupCon, and a replay of the cosine distillation on classification text.
- Cross-lingual: the multi-task recipe with rebalanced task weights and added English retrieval passages, to improve English-Turkish alignment.
- We weight-average the checkpoints described above into a single model.
Usage
sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("mrbesher/modernbert-tr-embed")
emb = model.encode(["bir cümle", "başka bir cümle"], normalize_embeddings=True)
q = model.encode(["soru"], prompt_name="query", normalize_embeddings=True)
d = model.encode(["döküman"], normalize_embeddings=True)The retrieval (see config_sentence_transformers.json): Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery:{text}.
ONNX Runtime
The onnx/ folder has the grapgh for the token embeddings, it includes 3 graphs: token embeddings, mean-pool and L2-norm.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("mrbesher/modernbert-tr-embed", backend="onnx",
model_kwargs={"file_name": "onnx/model_fp16.onnx"})Text Embeddings Inference (TEI)
TEI is L2-normalizes by default. Pass per-request prompts for retrieval queries.
text-embeddings-router --model-id mrbesher/modernbert-tr-embed --dtype float16Encoderfile
See the companion encoderfile repo.
Training data
Turkish retrieval (msmarco-tr, Squad-TR train/dev), Turkish NLI (boun-tabi/nli_tr train), Turkish STS-B (train), and Turkish classification-domain text (product reviews, news, social), all teacher-supervised. We check for leaks with text-hash against every MTEB(Turkish) test split.
License & attribution
- License:
apache-2.0.
