tuskanny/ms_marco_colbertv2
MS MARCO v1 Passage, ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries. Source Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/ms_marco_colbertv2.
MS MARCO v1 Passage, ColBERTv2
Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.
Source
- Collection: MS MARCO v1 passage (
ir_datasetsmsmarco-passage), 8,841,823 passages - Queries: dev/small, 6,980 queries and 7,437 qrels (
msmarco-passage/dev/small) - Document order: passage id order (row i is pid i)
Encoding
- Model: ColBERTv2 (
colbert-ir/colbertv2.0, BERT-base-uncased tokenizer) - Documents start with
[CLS]and the ColBERT[D]marker (token id 2) - Queries: always 32 vectors (
[MASK]expansion, no zero padding) - Vectors: 128-d, L2-normalized
- Not recorded: exact checkpoint revision, encoding library, document length cap (the longest document has 300 vectors), and whether punctuation was dropped
Statistics
Files
Document ids are MS MARCO passage ids and equal the row index: row i of doclens.npy is passage i.
documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16. In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).
Relevance judgments
The official MS MARCO passage dev/small qrels (qrels.dev.small.tsv): 7,437 relevant query-passage pairs over the 6,980 queries, all with relevance 1 (ir_datasets msmarco-passage/dev/small). Standard metric: MRR@10 (RR@10 in ir_measures).
