ColBERT
Datasets
All datasets matching “ColBERT”colbert-retrieval-mined-examples
Cantivia internal retrieval training artifacts
Proprietary. All rights reserved.
This repository stores working artifacts of Cantivia's reranker training pipeline: mined candidate
pools, hard negatives, group caches, relevance scores and their manifests. It is published only so
Cantivia's own jobs can reach it; it is not a public dataset, and nothing here is released for
download, redistribution, derivative works or model training.
Licence
Cantivia's own… See the full description on the dataset page: https://huggingface.co/datasets/Feargal/colbert-retrieval-mined-examples.lotte_passagesLoTTE Passages Dataset for ColBERTv2msmarco_answerai_colbert_small_embeddings
MS MARCO ColBERT Embeddings
Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1.
Dataset Structure
The dataset contains:
data/corpus/: 177 parquet files with document embeddings
data/queries/: 11 parquet files with query embeddings
data/qrels/train.parquet: Relevance judgments (532,751 pairs)
Usage
from datasets import load_dataset
# Load from directory (recommended for large datasets)
corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.lotteLoTTE Passages Dataset for ColBERTv2msmarco_token_score_colbertx_xlmr_large_zs_en_en
Dataset Card for "msmarco_token_score_colbertx_xlmr_large_zs_en_en"
More Information needed
hotpotqa_colbert
