Team Ai
Modelpublic

DataScience-UIBK/uibk-embedding-v1

sourceHugging Faceapache-2.0updated 7h agoView on Hugging Face
0likes17downloads
Model Card

<div align="center"> <img src="https://huggingface.co/DataScience-UIBK/uibk-embedding-v1/resolve/main/assets/logo.png" alt="UIBK-Embedding-v1" width="256">

![University of Innsbruck](https://www.uibk.ac.at/en/) ![Hugging Face](https://huggingface.co/DataScience-UIBK) ![BEIR](https://huggingface.co/spaces/mteb/leaderboard) ![License](https://www.apache.org/licenses/LICENSE-2.0)

</div>

<h1 align="center">UIBK-Embedding-v1</h1>

<h3 align="center">Unifying late Interaction and Bi-encoders with Knowledge distillation</h3>

<p align="center"> <a href="https://huggingface.co/DataScience-UIBK/uibk-embedding-v1">UIBK-Embedding-v1</a> | <a href="https://github.com/lightonai/pylate">PyLate</a> | <a href="https://github.com/lightonai/fast-plaid">FastPLAID</a> | <a href="https://github.com/embeddings-benchmark/mteb">MTEB</a> </p>


About UIBK-Embedding

Single-vector (dense) retrievers and multi-vector (late-interaction, ColBERT) retrievers are usually presented as a choice. They are strong in different places. A dense vector captures what a text is about, which is what counts for arguments, duplicate questions and scientific topics. Token-level matching finds the exact entity or phrase, which is what counts for multi-hop and factoid questions. We measured this on two fine-tunes of our own backbone: the dense one was ahead on ArguAna, SCIDOCS, FiQA and Touché, the ColBERT one on HotpotQA, NQ, MS MARCO and DBPedia, and simply adding their two scores beat the better of the two by more than two points on NanoBEIR.

The usual way to get both is to run two models, keep two indexes and fuse two result lists. UIBK-Embedding does it inside one model:

  • —one encoder pass produces the token vectors and the dense embedding,
  • —the dense embedding is stored as three extra vectors in the same list as the token vectors,
  • —a plain MaxSim over that list is the sum of the two scores.

Nothing changes for the index or the scorer: the model loads as a PyLate ColBERT model and works with PLAID as it is.

The name spells the recipe, Unifying late Interaction and Bi-encoders with Knowledge distillation, and UIBK is the University of Innsbruck, where the model was built.

UIBK-Embedding-v1

UIBK-Embedding-v1 is an English retrieval model built on ModernBERT-base (179M parameters) that holds a late-interaction (ColBERT) retriever and a dense retriever in one encoder and one index. It was trained by the Data Science group of the University of Innsbruck using PyLate.

Notably it:

  • —Reaches 59.79 average nDCG@10 on BEIR (15 datasets), 2.57 points above LateOn (57.22), and is ahead of it on all 15 datasets.
  • —Is a drop-in ColBERT model: its output is an ordinary list of 128-dimensional vectors, and plain MaxSim over it equals token MaxSim + 0.75 × dense cosine. No second index, no score fusion at query time.
  • —Reads long inputs: documents up to 2048 tokens and queries up to 256 tokens.
  • —Uses open data and an open teacher: LightOn's public pre-training and fine-tuning collections, and the public cross-encoder mxbai-rerank-large-v2 as the teacher.

Two things should be said for a fair reading of the comparison with LateOn. The model has 30M more parameters (its second 4-layer top), and it reads documents up to 2048 tokens, where LateOn's reported numbers use 300; with 2048-token documents LateOn scores 57.69 in our evaluation. And unlike LateOn it was trained with knowledge distillation and with prompts, the two steps LightOn name as the next ones for their own models.

Results

All numbers of UIBK-Embedding-v1 below come from MTEB 2.21.6 with its stock PyLate wrapper and a PLAID index with default settings.

BEIR (15 datasets, NDCG@10)

ModelAverageSize (M)Embed dimArguAnaCQADupstackRetrievalClimateFEVERDBPediaFEVERFiQA2018HotpotQAMSMARCONFCorpusNQQuoraRetrievalSCIDOCSSciFactTRECCOVIDTouche2020
ColBERTv248.6311012846.5038.3017.6045.2078.5035.4067.5046.0033.7052.4085.5015.4068.9072.6026.00
Jina-ColBERT-v251.8560012836.6040.8023.9047.1080.5040.8076.6046.9034.6064.0088.7018.6067.8083.4027.40
ColBERT-small53.79339650.0938.7533.0745.5890.9641.1576.1143.5037.3059.1087.7218.4274.7784.5925.69
GTE-ModernColBERT-v154.7514912847.5241.0831.3347.5687.6745.2577.4845.6037.8361.6286.7119.2276.3384.8431.25
ColBERT-Zero55.3914912852.8241.4135.9047.4390.5242.5079.4545.9537.2161.8285.1919.8476.3378.2736.24
LateOn-unsupervised50.1114912843.1247.7118.7643.3665.7451.9468.1737.5137.1558.4189.4821.1376.8969.8122.53
LateOn57.2214912850.5247.3639.6745.9992.0253.1279.9845.6737.7963.9189.6721.9076.6183.6030.52
LateOn, 2048-token documents57.6914912850.1947.1740.1346.2293.0754.1679.9745.8337.9363.2889.6621.8677.2083.6235.09
UIBK-Embedding-v159.7917912856.5648.5544.8949.2293.2853.7381.5046.8238.7967.8089.9223.6078.5783.8239.78

The numbers of the other models are those of the LateOn model card. The row "LateOn, 2048-token documents" is our own evaluation of LateOn with the document length raised from 300 to 2048.

UIBK-Embedding-v1 reaches 59.79 average NDCG@10, 2.57 points above LateOn, and is ahead of it on all 15 datasets. The largest gains are on Touche2020 (+9.26), ArguAna (+6.04), ClimateFEVER (+5.22), NQ (+3.89) and DBPedia (+3.23). These are datasets of both kinds: ArguAna and Touché are the long-query, argument-style datasets where dense models lead, NQ and DBPedia are entity-centred datasets where late interaction leads. The model does not trade one family's datasets for the other's, which is what the two heads are for.

Long documents explain part of the gap, not most of it: LateOn gains 0.47 on average from 2048-token documents (almost all of it on Touché), which leaves about two points to the model and its training.

Like LateOn and the other models trained on this collection, the model is not zero-shot on the BEIR datasets whose training splits are in the fine-tuning data (MS MARCO, NQ, HotpotQA, FEVER, FiQA). Training queries that equal a BEIR test or development query were removed from every training set (see Training Details).

NanoBEIR (13 datasets, NDCG@10)

ModelAverageNanoArguAnaNanoClimateFeverNanoDBPediaNanoFEVERNanoFiQA2018NanoHotpotQANanoMSMARCONanoNFCorpusNanoNQNanoQuoraNanoSCIDOCSNanoSciFactNanoTouche2020
UIBK-Embedding-v170.5559.8451.6569.2096.5764.3492.7268.9939.2282.1496.9945.4083.9366.14

LateOn's card has no NanoBEIR numbers. In our development evaluation of NanoBEIR (exact MaxSim search, no index) LateOn scores 67.50 and UIBK-Embedding-v1 70.43.

How It Works

<div align="center"> <img src="https://huggingface.co/DataScience-UIBK/uibk-embedding-v1/resolve/main/assets/architecture.png" alt="Architecture of UIBK-Embedding-v1" width="1024"> </div>

One trunk, two tops. The ModernBERT embeddings and the first 18 layers are shared. Each head then has its own copy of the last 4 layers, so the token head and the dense head are free to make different errors. A single encoder with all 22 layers shared reached a plateau that three different training signals could not move; the two tops moved it.

Token head. The usual ColBERT projection (768 → 128), one vector per token. For long queries only the first 64 tokens take part in token matching.

Dense head. Latent-attention pooling (as in NV-Embed) over the dense top's states gives one 768-dimensional sentence vector, which reads the whole query or document. A 768-dimensional vector does not fit into a 128-dimensional slot, so it is rotated and cut into 3 slices of 127 dimensions: the 3 slot vectors. They take the first three positions of the output list.

One MaxSim. Every output vector has 128 dimensions and unit norm. Dimension 0 is a tag, positive for token vectors and negative for slot vectors:

token vector    [ +sqrt(0.1)   ; sqrt(0.9)   * u_t ]     u_t: the projected token state
slot vector s   [ -sqrt(0.775) ; sqrt(0.225) * g_s ]     g_s: slice s of the rotated dense vector, normalised

token x token   =  0.1   + 0.9   * cos(u, u')
slot s x slot s =  0.775 + 0.225 * cos(g_s, g'_s)
token x slot    = -0.28  + (about 0)                      never the maximum

Because of the tag, a query token only ever picks a document token and a slot only a slot. Plain MaxSim is therefore, up to a constant per query,

MaxSim(q, d)  =  0.9 * [ token MaxSim  +  0.75 * mean slice cosine ]

the fixed-weight sum of a ColBERT score and a dense score, computed by any ColBERT index and scorer with no change.

How We Got Here

<div align="center"> <img src="https://huggingface.co/DataScience-UIBK/uibk-embedding-v1/resolve/main/assets/path.png" alt="BEIR average after each change, from our dense model to UIBK-Embedding-v1" width="1024"> </div>

Every number in this section is a full BEIR run (15 datasets) of that model in our development evaluation: MTEB with a multi-GPU wrapper. The released model scores 59.74 there and 59.79 with MTEB's stock wrapper.

  1. 1.A dense model of our own (56.59). We started from raw ModernBERT-base and pre-trained it contrastively on LightOn's public pre-training collection, with latent-attention pooling in place of [CLS] pooling (in a controlled ablation it gave +0.8 NanoBEIR at an identical training loss). After fine-tuning with hard negatives and cross-encoder distillation, the average of four fine-tunes reached 56.59: below LateOn, and behind it mostly on HotpotQA (−6.5), MS MARCO (−3.7) and NQ (−3.6), which hold two thirds of all the points it lost.
  2. 2.The two families fail in opposite places. A ColBERT fine-tune of the same backbone won those datasets back and lost the dense model's lead elsewhere. The sum of the two models' scores reached 69.93 on NanoBEIR against 67.47 and 66.61 for its parts. The signals are complementary, so the question became how to get both from one model.
  3. 3.One encoder, one index (57.22). Both heads on one encoder, the dense vector written into the token list so that MaxSim adds the two scores.
  4. 4.Two tops (57.61) and long documents (58.14). Each head got its own copy of the last 4 layers (+30M parameters), and documents are read up to 2048 tokens at inference.
  5. 5.Data (58.61). 412k mined pairs from the sources where the model was weak (Quora, StackExchange, S2ORC), 151k long-query pairs with a query length of 256 tokens, and hard negatives re-mined with our own dense model and scored by the cross-encoder.
  6. 6.Distil the score that is evaluated (59.08, then 59.20). The largest training gain: the cross-encoder's KL loss is applied to the fused score, the one MaxSim computes at inference, not only to each head. This was the first model ahead of LateOn on all 15 datasets. Using all 14 negatives per query instead of 10 added a little more.
  7. 7.Learning rate (59.74). The last lever, and a simple one. The recipe had been tuned at 1e-5. Raising it step by step gave 59.25 and 59.47 (2e-5, two seeds), 59.62 and 59.45 (3e-5, two seeds), 59.69 (4e-5), 59.62 (5e-5) and 59.74 (7e-5), the released model.

What did not help. Each of these was a full training run with a full BEIR evaluation:

  • —a larger dense model (435M parameters, 58.97 on BEIR by itself) as a target for the dense head or as a second score teacher: 58.41 to 59.02;
  • —a second teacher next to the cross-encoder, a 9B-parameter late-interaction model: 59.14 to 59.46, against 59.20 to 59.47 without it;
  • —softer distillation targets: 58.14 and 59.04;
  • —a second epoch: 59.44; averaging the weights of two seeds: 59.38, the mean of its parts;
  • —rank-based distillation losses and masking of the negatives the teacher prefers to the positive: 58.79 to 59.18;
  • —dropping the feature anchors to the earlier single-head models: 58.96.

Two seeds of one recipe differ by about 0.2, so differences below that are noise. Development decisions were taken on BEIR itself, which makes the last tenths optimistic; see Limitations.

Model Details

Model Description

  • —Model Type: PyLate ColBERT model with an additional dense head
  • —Base model: ModernBERT-base, with our own contrastive pre-training
  • —Parameters: 179M (shared trunk and token top 149.0M, dense top 20.1M, output layer 9.9M)
  • —Document Length: 2048 tokens
  • —Query Length: 256 tokens (the token head uses the first 64, the dense head all of them)
  • —Output Dimensionality: 128 per vector; one vector per token, of which the first 3 carry the dense embedding
  • —Similarity Function: MaxSim
  • —Prompts: query: and document: (required)
  • —Language: English
  • —License: Apache 2.0

Model Sources

Full Model Architecture

ColBERT(
  (0): Transformer({'max_seq_length': 299, 'do_lower_case': False, 'shared_layers': 18, 'architecture': 'ModernBertModel'})
  (1): HybridTokens({'dim': 768, 'token_dim': 128, 'slots': 3, 'local_weight': 0.9, 'slot_weight': 0.225, 'num_latents': 512, 'num_heads': 8, 'ff_mult': 4, 'query_marker_id': 50368, 'query_local_cap': 64, 'query_local_ref': 0, 'query_local_power': 1.0, 'query_slot_weight': -1.0, 'doc_slot_weight': -1.0, 'slot_copies': 1}, params=9.9M)
)

The input lengths are set by PyLate at encoding time from the model's configuration: 256 tokens for queries and 2048 for documents.

Usage

The model is used with PyLate. Two things differ from an ordinary ColBERT model:

  • —it needs trust_remote_code=True, because two small modules in this repository define the two-top encoder and the output layer;
  • —the prompts must be passed when encoding: prompt_name="query" for queries and prompt_name="document" for documents.

First install the PyLate library:

bash
pip install -U pylate "transformers>=5.3" "sentence-transformers>=5.3"

Quick start

python
from pylate import models, rank

model = models.ColBERT("DataScience-UIBK/uibk-embedding-v1", trust_remote_code=True)

query = "Which planet is known as the Red Planet?"
documents = [
    "Venus is often called Earth's twin because of its similar size and proximity.",
    "Mars, known for its reddish appearance, is often referred to as the Red Planet.",
    "Jupiter, the largest planet in our solar system, has a prominent red spot.",
    "Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
]

queries_embeddings = model.encode([query], is_query=True, prompt_name="query")
documents_embeddings = model.encode(documents, is_query=False, prompt_name="document")
print(queries_embeddings[0].shape, documents_embeddings[0].shape)
# (14, 128) (18, 128)

# MaxSim late-interaction scoring (higher is more relevant)
ranking = rank.rerank(
    documents_ids=[[1, 2, 3, 4]],
    queries_embeddings=queries_embeddings,
    documents_embeddings=[documents_embeddings],
)
print([(d["id"], round(d["score"], 2)) for d in ranking[0]])
# [(2, 13.45), (4, 13.36), (3, 13.26), (1, 13.05)]

Retrieval

Use this model with PyLate to index and retrieve documents. The index uses FastPLAID for efficient similarity search.

Indexing documents

Load the ColBERT model and initialize the PLAID index, then encode and index your documents:

python
from pylate import indexes, models, retrieve

# Step 1: Load the ColBERT model
model = models.ColBERT(
    model_name_or_path="DataScience-UIBK/uibk-embedding-v1",
    trust_remote_code=True,
)

# Step 2: Initialize the PLAID index
index = indexes.PLAID(
    index_folder="pylate-index",
    index_name="index",
    override=True,  # This overwrites the existing index if any
)

# Step 3: Encode the documents
documents_ids = ["1", "2", "3"]
documents = ["document 1 text", "document 2 text", "document 3 text"]

documents_embeddings = model.encode(
    documents,
    batch_size=32,
    is_query=False,  # Ensure that it is set to False to indicate that these are documents, not queries
    prompt_name="document",  # The document prompt is required
    show_progress_bar=True,
)

# Step 4: Add document embeddings to the index by providing embeddings and corresponding ids
index.add_documents(
    documents_ids=documents_ids,
    documents_embeddings=documents_embeddings,
)

Note that you do not have to recreate the index and encode the documents every time. Once you have created an index and added the documents, you can re-use the index later by loading it:

python
# To load an index, simply instantiate it with the correct folder/name and without overriding it
index = indexes.PLAID(
    index_folder="pylate-index",
    index_name="index",
)
Retrieving top-k documents for queries

Once the documents are indexed, you can retrieve the top-k most relevant documents for a given set of queries. To do so, initialize the ColBERT retriever with the index you want to search in, encode the queries and then retrieve the top-k documents to get the top matches ids and relevance scores:

python
# Step 1: Initialize the ColBERT retriever
retriever = retrieve.ColBERT(index=index)

# Step 2: Encode the queries
queries_embeddings = model.encode(
    ["query for document 3", "query for document 1"],
    batch_size=32,
    is_query=True,  # Ensure that it is set to True to indicate that these are queries
    prompt_name="query",  # The query prompt is required
    show_progress_bar=True,
)

# Step 3: Retrieve top-k documents
scores = retriever.retrieve(
    queries_embeddings=queries_embeddings,
    k=10,  # Retrieve the top 10 matches for each query
)

For large collections of long documents, pass batch_size="auto" to indexes.PLAID(...): with 2048-token documents the fixed default scoring batch can run out of GPU memory.

Reranking

If you only want to use the model to perform reranking on top of your first-stage retrieval pipeline without building an index, you can simply use rank function and pass the queries and documents to rerank:

python
from pylate import rank, models

queries = [
    "query A",
    "query B",
]

documents = [
    ["document A", "document B"],
    ["document 1", "document C", "document B"],
]

documents_ids = [
    [1, 2],
    [1, 3, 2],
]

model = models.ColBERT(
    model_name_or_path="DataScience-UIBK/uibk-embedding-v1",
    trust_remote_code=True,
)

queries_embeddings = model.encode(
    queries,
    is_query=True,
    prompt_name="query",
)

documents_embeddings = model.encode(
    documents,
    is_query=False,
    prompt_name="document",
)

reranked_documents = rank.rerank(
    documents_ids=documents_ids,
    queries_embeddings=queries_embeddings,
    documents_embeddings=documents_embeddings,
)

Training Details

Training Stages

  1. 1.Pre-training. ModernBERT-base with latent-attention pooling, contrastive loss with in-batch negatives (temperature 0.02), 33,000 steps with a batch of 16,384 pairs, on a 331M-pair sample of LightOn's curated pre-training collection, prompts query: and document: .
  2. 2.Single-head models. A dense model (the average of four fine-tunes) and a ColBERT model, both fine-tuned from the pre-trained backbone with hard negatives. They serve as feature anchors in the last stage.
  3. 3.UIBK-Embedding-v1. Both heads on the two-top encoder, one epoch (16,597 steps, batch 128, learning rate 7e-5 with 5% warm-up, 14 hard negatives per query, queries up to 256 tokens, documents up to 512 tokens), about 4.3 hours on 4 H100 GPUs. The loss is the sum of
  4. 4.a contrastive loss on each head (temperature 0.02),
  5. 5.a KL loss to the cross-encoder's scores on the dense head and on the fused score,
  6. 6.feature anchors that keep each head close to the corresponding single-head model.

Training Data

  • —LightOn's public English fine-tuning collection (MS MARCO, NQ, HotpotQA, FEVER, FiQA, SQuAD v2, TriviaQA, MIRACL) with its mined negatives; 4 more negatives per query were mined with our dense model.
  • —About 560k pairs mined from the pre-training collection (Quora, StackExchange, S2ORC, Reddit), 151k of them with long queries.
  • —Teacher scores from mxbai-rerank-large-v2.
  • —Overlap with BEIR. Rows whose query equals a BEIR test or development query were removed: 37,391 rows of the pre-training sample and 228 rows of the fine-tuning collection. For the mined pairs, a pair was removed if its query or its document equals a BEIR query.
  • —For transparency: three of the four dense fine-tunes behind the dense anchor were trained with an additional ranking loss to our fine-tune of stella_en_400M_v5, the 435M model mentioned above.

Framework Versions

  • —Python: 3.12
  • —Sentence Transformers: 5.3.0
  • —PyLate: 1.6.0
  • —Transformers: 5.3.0
  • —PyTorch: 2.9.0+cu128
  • —Accelerate: 1.15.0
  • —Datasets: 3.6.0
  • —Tokenizers: 0.22.2

Limitations

  • —English only.
  • —Multi-vector output: a ColBERT index is larger than a single-vector index.
  • —Needs trust_remote_code=True, transformers >= 5.3 and the two prompts. The MultiVectorEncoder class of Sentence Transformers 6 has not been tested with this model.
  • —Development decisions (data, losses, learning rate) were taken on BEIR itself, and the model has not yet been evaluated on a decontaminated or held-out benchmark. Read the BEIR average with that in mind.

Citation

BibTeX

UIBK-Embedding-v1
bibtex
@misc{uibk2026embedding,
  title={UIBK-Embedding-v1: Unifying late Interaction and Bi-encoders with Knowledge distillation},
  author={{Data Science Group, University of Innsbruck}},
  year={2026},
  howpublished={\url{https://huggingface.co/DataScience-UIBK/uibk-embedding-v1}},
}
DenseOn and LateOn
bibtex
@misc{sourty2026denseonlateon,
  title={DenseOn with the LateOn: Open State-of-the-Art Single and Multi-Vector Models},
  author={Sourty, Raphael and Chaffin, Antoine and Weller, Orion and Demoura, Paulo and Chatelain, Amelie},
  year={2026},
  howpublished={\url{https://huggingface.co/blog/lightonai/denseon-lateon}},
}
PyLate
bibtex
@inproceedings{DBLP:conf/cikm/ChaffinS25,
  author       = {Antoine Chaffin and
                  Rapha{\"{e}}l Sourty},
  editor       = {Meeyoung Cha and
                  Chanyoung Park and
                  Noseong Park and
                  Carl Yang and
                  Senjuti Basu Roy and
                  Jessie Li and
                  Jaap Kamps and
                  Kijung Shin and
                  Bryan Hooi and
                  Lifang He},
  title        = {PyLate: Flexible Training and Retrieval for Late Interaction Models},
  booktitle    = {Proceedings of the 34th {ACM} International Conference on Information
                  and Knowledge Management, {CIKM} 2025, Seoul, Republic of Korea, November
                  10-14, 2025},
  pages        = {6334--6339},
  publisher    = {{ACM}},
  year         = {2025},
  url          = {https://github.com/lightonai/pylate},
  doi          = {10.1145/3746252.3761608},
}
Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084"
}

Acknowledgements

This model stands on open work. We thank LightOn for PyLate, FastPLAID, the public pre-training and fine-tuning collections and LateOn, the reference that set the bar; Answer.AI and LightOn for ModernBERT; Mixedbread for the cross-encoder used as the teacher; the authors of NV-Embed for latent-attention pooling; and the teams behind Sentence Transformers, MTEB and BEIR.

About the Logo

Innsbruck means "bridge over the Inn". The logo is that bridge. The white arch stones are the token vectors of the late-interaction head, and the three golden keystones that hold the arch together are the three slot vectors of the dense head, gold like the roof of the Goldenes Dachl. Through the arch you see the Nordkette, below it runs the Inn. One arch, two kinds of stones: one model, two retrievers.