datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco_answerai_colbert_small_embeddings
MS MARCO ColBERT Embeddings
Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1.
Dataset Structure
The dataset contains:
data/corpus/: 177 parquet files with document embeddings
data/queries/: 11 parquet files with query embeddings
data/qrels/train.parquet: Relevance judgments (532,751 pairs)
Usage
from datasets import load_dataset
# Load from directory (recommended for large datasets)
corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.pgturbohybrid_dbpedia_colbert
johannhartmann/pgturbohybrid_dbpedia_colbert
Precomputed DBpedia ColBERT multivectors for pgturbohybrid benchmark runs.
The dataset stores packed little-endian float16 values for the document and query embeddings.
Importing these rows into PostgreSQL avoids llama.cpp embedding generation
during retrieval/index benchmarks.
This compact export stores half-precision values and reconstructs turbohybrid_multivector values on import. Use it for benchmark loading where avoiding runtime… See the full description on the dataset page: https://huggingface.co/datasets/johannhartmann/pgturbohybrid_dbpedia_colbert.fiqa_modern_colbert
FiQA-2018, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding
57,638 documents, 648 queries, 1,706 qrels
Text given to the encoder for each document: the passage text (FiQA documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_modern_colbert.scidocs_modern_colbert
SCIDOCS, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
25,657 documents, 1,000 queries, 29,928 qrels
Text given to the encoder for each document: title + " " + text (BEIR… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_modern_colbert.scidocs_colbertv2
SCIDOCS, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
25,657 documents, 1,000 queries, 29,928 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_colbertv2.scifact_answerai_colbert_small
scifact_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_answerai_colbert_small.fiqa_answerai_colbert_small
fiqa_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR fiqa (beir/fiqa/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/fiqa/test (ir_datasets 0.6.3), which downloads fiqa.zip (md5 17918ed23cd04fb15047f73e6c3bd9d9). BEIR also publishes this corpus on the Hub as BeIR/fiqa, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from… See the full description on the dataset page: https://huggingface.co/datasets/robro612/fiqa_answerai_colbert_small.nfcorpus_colbertv2
NFCorpus, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding
3,633 documents, 323 queries, 12,334 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_colbertv2.nfcorpus_answerai_colbert_small
nfcorpus_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_answerai_colbert_small.scidocs_answerai_colbert_small
scidocs_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scidocs_answerai_colbert_small.lotte_pooled_dev_search_answerai_colbert_small
lotte_pooled_dev_search_answerai_colbert_small
Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_answerai_colbert_small.msmarco_answerai_colbert_small
msmarco_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR msmarco (beir/msmarco/dev), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/msmarco/dev (ir_datasets 0.6.3), which downloads msmarco.zip (md5 444067daf65d982533ea17ebd59501e4). BEIR also publishes this corpus on the Hub as BeIR/msmarco, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/msmarco_answerai_colbert_small.ms_marco_colbertv2
MS MARCO v1 Passage, ColBERTv2
Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.
Source
Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages
Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small)
Document order: passage id order (row i is pid i)
Encoding
Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/ms_marco_colbertv2.trec-covid_answerai_colbert_small
trec-covid_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_answerai_colbert_small.fiqa_colbertv2
FiQA-2018, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding
57,638 documents, 648 queries, 1,706 qrels
Text given to the encoder for each document: the passage text (FiQA documents have no title). The… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_colbertv2.nfcorpus_modern_colbert
NFCorpus, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding
3,633 documents, 323 queries, 12,334 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_modern_colbert.lotte_pooled_colbertv2
LoTTE pooled (dev), ColBERTv2
Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries.
Source
Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages
Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels)
Document order: passage id order (row i is pid i)
Encoding
Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
Documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/lotte_pooled_colbertv2.clini-colbert-pairs-dev-v2NQ-colbert-10k-case-entity
