datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
colbert-retrieval-mined-examples
Cantivia internal retrieval training artifacts
Proprietary. All rights reserved.
This repository stores working artifacts of Cantivia's reranker training pipeline: mined candidate
pools, hard negatives, group caches, relevance scores and their manifests. It is published only so
Cantivia's own jobs can reach it; it is not a public dataset, and nothing here is released for
download, redistribution, derivative works or model training.
Licence
Cantivia's own… See the full description on the dataset page: https://huggingface.co/datasets/Feargal/colbert-retrieval-mined-examples.lotte_passagesLoTTE Passages Dataset for ColBERTv2lotteLoTTE Passages Dataset for ColBERTv2msmarco_answerai_colbert_small_embeddings
MS MARCO ColBERT Embeddings
Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1.
Dataset Structure
The dataset contains:
data/corpus/: 177 parquet files with document embeddings
data/queries/: 11 parquet files with query embeddings
data/qrels/train.parquet: Relevance judgments (532,751 pairs)
Usage
from datasets import load_dataset
# Load from directory (recommended for large datasets)
corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.msmarco_token_score_colbertx_xlmr_large_zs_en_en
Dataset Card for "msmarco_token_score_colbertx_xlmr_large_zs_en_en"
More Information needed
hotpotqa_colbertcolbert-wiki2017nq_colbertfiqa_modern_colbert
FiQA-2018, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding
57,638 documents, 648 queries, 1,706 qrels
Text given to the encoder for each document: the passage text (FiQA documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_modern_colbert.scidocs_modern_colbert
SCIDOCS, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
25,657 documents, 1,000 queries, 29,928 qrels
Text given to the encoder for each document: title + " " + text (BEIR… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_modern_colbert.scidocs_colbertv2
SCIDOCS, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
25,657 documents, 1,000 queries, 29,928 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_colbertv2.scifact_answerai_colbert_small
scifact_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_answerai_colbert_small.msmarco-v2-colbertv2-fp32
MS MARCO v2 — ColBERTv2 fp32 Multi-Vector Embeddings (uncompressed)
Per-token multi-vector embeddings for the full MS MARCO v2 passage corpus
(~138.4M passages) plus the dev / dev2 queries, produced with the official
ColBERTv2 checkpoint (colbert-ir/colbertv2.0).
Precision: fp32 (NO residual quantization, NO pooling)
Dim: 128 per token
Corpus token vectors: ~9.41B (avg ~68 tokens/passage)
Total corpus size: ~4.82 TB
These embeddings are uncompressed on purpose (research on… See the full description on the dataset page: https://huggingface.co/datasets/yaooooo233/msmarco-v2-colbertv2-fp32.xpr_colbertnq_colbertv2nfcorpus_colbertv2
NFCorpus, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding
3,633 documents, 323 queries, 12,334 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_colbertv2.msmarco_psg_v1.colbertv2
msmarco_psg_v1.colbertv2
Description
This is a ColBERT v2 index created via pyterrier_colbert2. This is the index that is used in the SIGIR 2026 PLAID-PRF paper.
Usage
# Download and load the index from HuggingFace
import pyterrier as pt
index = pt.Artifact.from_hf('pyterrier/msmarco_psg_v1.colbertv2',
plaid_mode=True, ncells=4,
centroid_score_threshold=0.4, ndocs=4096)
# TREC-DL 2019 pt.Experiment:
from pyterrier_colbert.ranking import _prf… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/msmarco_psg_v1.colbertv2.ColBERT_Humor_Detection
ColBERT_Humor
Dataset Summary
ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.ms_marco_colbertv2
MS MARCO v1 Passage, ColBERTv2
Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.
Source
Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages
Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small)
Document order: passage id order (row i is pid i)
Encoding
Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/ms_marco_colbertv2.fiqa_answerai_colbert_small
fiqa_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR fiqa (beir/fiqa/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/fiqa/test (ir_datasets 0.6.3), which downloads fiqa.zip (md5 17918ed23cd04fb15047f73e6c3bd9d9). BEIR also publishes this corpus on the Hub as BeIR/fiqa, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from… See the full description on the dataset page: https://huggingface.co/datasets/robro612/fiqa_answerai_colbert_small.trec-covid_answerai_colbert_small
trec-covid_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_answerai_colbert_small.lotte_pooled_dev_search_answerai_colbert_small
lotte_pooled_dev_search_answerai_colbert_small
Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_answerai_colbert_small.nfcorpus_answerai_colbert_small
nfcorpus_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_answerai_colbert_small.scidocs_answerai_colbert_small
scidocs_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scidocs_answerai_colbert_small.msmarco_answerai_colbert_small
msmarco_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR msmarco (beir/msmarco/dev), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/msmarco/dev (ir_datasets 0.6.3), which downloads msmarco.zip (md5 444067daf65d982533ea17ebd59501e4). BEIR also publishes this corpus on the Hub as BeIR/msmarco, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/msmarco_answerai_colbert_small.fiqa_colbertv2
FiQA-2018, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding
57,638 documents, 648 queries, 1,706 qrels
Text given to the encoder for each document: the passage text (FiQA documents have no title). The… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_colbertv2.lotte_pooled_colbertv2
LoTTE pooled (dev), ColBERTv2
Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries.
Source
Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages
Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels)
Document order: passage id order (row i is pid i)
Encoding
Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
Documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/lotte_pooled_colbertv2.nfcorpus_modern_colbert
NFCorpus, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding
3,633 documents, 323 queries, 12,334 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_modern_colbert.msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__
msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__
Description
This is the PyTerrier JPQIndex for MSMARCO v1 passage corpus, which corresponds to an result from the SIGIR 2026 reproducibility paper.
Usage
# Load the artifact
import pyterrier as pt
import pyterrier_dr
index = pt.Artifact.from_hf('jpq-repro/msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__')
model =… See the full description on the dataset page: https://huggingface.co/datasets/jpq-repro/msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__.msmarco-train-distil-colbert-v2
