robro612/nfcorpus_modernbert_xtr
nfcorpus_modernbert_xtr Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with robro612/ModernBERT-XTR at revision 8c06fce0b8d2387be98183582eb9600af7bd1b8a. Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded through ir_datasets, not… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_modernbert_xtr.
nfcorpusmodernbertxtr
Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with [robro612/ModernBERT-XTR](https://huggingface.co/robro612/ModernBERT-XTR) at revision 8c06fce0b8d2387be98183582eb9600af7bd1b8a.
Source data: ir_datasets beir/nfcorpus/test (irdatasets 0.6.3), which downloads [nfcorpus.zip](https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/nfcorpus.zip) (md5 `a89dba18a62ef92f7d323ec890a0d38d`). BEIR also publishes this corpus on the Hub as [`BeIR/nfcorpus`](https://huggingface.co/datasets/BeIR/nfcorpus), whose card gives this dataset's license; the data here was loaded through irdatasets, not from that repo. Document, query and qrel ids are the source's own ids, unchanged.
Every document is one variable-length set of 128-d vectors; every query is one variable-length set of 128-d vectors. Documents and queries are stored at different precisions (fp16 and fp32 respectively), see Encoding.
Files
All positional indices (the gt_top*.tsv files, and the row order of every .npy file) refer to the order of doc_ids.npy and queries_ids.npy. Reordering either file invalidates the ground truth.
Statistics
Encoding
Ground truth: gt_top1000.tsv and gt_top100.tsv
Exact brute-force MaxSim top-1000 per query over the full corpus, from the vectors in this repo. gt_top100.tsv holds the first 100 ranks per query of the same lists (the original layout of these exports).
No header; tab-separated qidx docidx rank score:
qidx: 0-based row intoqueries_ids.npy/queries.npydocidx: 0-based position intodoc_ids.npy/doclens.npyrank: 1-based, descending scorescore:sum over the query's query_lens[qidx] vectors of max over the document's vectors of the dot product, computed in fp32 with the fp16 document vectors upcast to fp32. Expansion vectors are included in the sum. Printed to 6 decimals.
Retrieval quality
Sanity check of the vectors, not a leaderboard number: gt_top1000.tsv (exact MaxSim over the full corpus) scored against qrels.test.tsv with ir_measures.
Loading
import numpy as np
documents = np.load("documents.npy", mmap_mode="r") # [n_tokens, 128] float16
doclens = np.load("doclens.npy") # [n_docs] int32
offsets = np.concatenate([[0], np.cumsum(doclens)])
doc_ids = np.load("doc_ids.npy") # [n_docs] str
def document(i):
return documents[offsets[i]:offsets[i + 1]] # [doclens[i], 128]
queries = np.load("queries.npy") # [n_queries, 32, 128] float32
query_lens = np.load("query_lens.npy") # [n_queries] int32
query_ids = np.load("queries_ids.npy") # [n_queries] str
def query(j):
return queries[j, :query_lens[j]] # [query_lens[j], 128]
def maxsim(q, d):
return (q @ d.astype(np.float32).T).max(axis=1).sum()Validation
Checks run by the exporter on the files exactly as written here:
- ✅ file set — missing=[] extra=[]
- ✅ documents.npy dtype/shape — <f2 (1047014, 128)
- ✅ doclens.npy dtype/shape — <i4 (3633,)
- ✅ doc_ids.npy is a string array — <U8 (3633,)
- ✅ queries.npy dtype/shape — <f4 (323, 32, 128)
- ✅ query_lens.npy dtype/shape — <i4 (323,)
- ✅ queries_ids.npy is a string array — <U10 (323,)
- ✅ sum(doclens) == n_tokens — 1047014 vs 1047014
- ✅ no empty documents — min doclen 24
- ✅ len(doc_ids) == len(doclens) == corpus size — 3633, 3633, 3633
- ✅ doc_ids unique
- ✅ query arrays aligned — 323, 323, 323
- ✅ doc and query dim agree — 128 / 128
- ✅ token_ids.npy dtype/shape — <u4 (1047014,)
- ✅ document vectors unit-norm (100k sample) — norm range [0.9995, 1.0005]
- ✅ query vectors unit-norm — norm range [1.000000, 1.000000]
- ✅ all vectors finite
- ✅ gt_top100.tsv has k rows per query — 32300 rows, k=100
- ✅ gt_top100.tsv rows grouped by qidx with ranks 1..k and descending scores
- ✅ gt_top100.tsv indices in range
- ✅ gt_top1000.tsv has k rows per query — 323000 rows, k=1000
- ✅ gt_top1000.tsv rows grouped by qidx with ranks 1..k and descending scores
- ✅ gt_top1000.tsv indices in range
- ✅ gttop100.tsv is the first 100 ranks of gttop1000.tsv
