Team Ai
Datasetpublic

tuskanny/ms_marco_colbertv2

MS MARCO v1 Passage, ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries. Source Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/ms_marco_colbertv2.

sourceHugging Faceupdated 17d agoView on Hugging Face
0likes88downloads
Dataset Card

MS MARCO v1 Passage, ColBERTv2

Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.

Source

  • —Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages
  • —Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small)
  • —Document order: passage id order (row i is pid i)

Encoding

  • —Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
  • —Documents start with [CLS] and the ColBERT [D] marker (token id 2)
  • —Queries: always 32 vectors ([MASK] expansion, no zero padding)
  • —Vectors: 128-d, L2-normalized
  • —Not recorded: exact checkpoint revision, encoding library, document length cap (the longest document has 300 vectors), and whether punctuation was dropped

Statistics

Token vectors (N)597,909,919
Avg vectors per document67.6 (min 4, max 300)
Vectors per query32

Files

FiledtypeShapeContent
documents.npyuint16 (<u2)[597909919, 128]Raw float16 bit patterns stored as uint16. Read with .view(np.float16)
doclens.npyint32[8841823]Vectors per document; sum == N
token_ids_per_token.npyint64[597909919]Input token id of each row of documents.npy
doc_ids.npyint64[8841823]MS MARCO pid (equals the row index)
queries.npyfloat32[6980, 32, 128]Query vectors
queries_ids.npyint64[6980]MS MARCO qid of each query

Document ids are MS MARCO passage ids and equal the row index: row i of doclens.npy is passage i.

documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16. In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).

Relevance judgments

The official MS MARCO passage dev/small qrels (qrels.dev.small.tsv): 7,437 relevant query-passage pairs over the 6,980 queries, all with relevance 1 (ir_datasets msmarco-passage/dev/small). Standard metric: MRR@10 (RR@10 in ir_measures).