datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lotteLoTTE Passages Dataset for ColBERTv2msmarco_answerai_colbert_small_embeddings
MS MARCO ColBERT Embeddings
Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1.
Dataset Structure
The dataset contains:
data/corpus/: 177 parquet files with document embeddings
data/queries/: 11 parquet files with query embeddings
data/qrels/train.parquet: Relevance judgments (532,751 pairs)
Usage
from datasets import load_dataset
# Load from directory (recommended for large datasets)
corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.pgturbohybrid_dbpedia_colbert
johannhartmann/pgturbohybrid_dbpedia_colbert
Precomputed DBpedia ColBERT multivectors for pgturbohybrid benchmark runs.
The dataset stores packed little-endian float16 values for the document and query embeddings.
Importing these rows into PostgreSQL avoids llama.cpp embedding generation
during retrieval/index benchmarks.
This compact export stores half-precision values and reconstructs turbohybrid_multivector values on import. Use it for benchmark loading where avoiding runtime… See the full description on the dataset page: https://huggingface.co/datasets/johannhartmann/pgturbohybrid_dbpedia_colbert.xpr_colbertnanobeir-colbert-token-vectors
NanoBEIR ColBERT token vectors
Per-token ColBERT vectors for all 13 NanoBEIR datasets, from two late-interaction models:
folder
model
revision
encoder dtype
lateon/
lightonai/LateOn
62911e105059585d244384c7d17826e35f669c17
fp32
iso/
topk-io/Iso-ModernColBERT
b95d9608ad424fdd7bfd045001576f59cbb89a98
bf16
pplx-late-0.6b/
perplexity-ai/pplx-embed-v2-late-0.6b
dd4e95b836a73f6f0c32e46ea127c0b86b02169e
bf16
Each model covers 56,723 documents (7,585,402 kept… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/nanobeir-colbert-token-vectors.scidocs_modern_colbert
SCIDOCS, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
25,657 documents, 1,000 queries, 29,928 qrels
Text given to the encoder for each document: title + " " + text (BEIR… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_modern_colbert.scidocs_colbertv2
SCIDOCS, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
25,657 documents, 1,000 queries, 29,928 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_colbertv2.scifact_answerai_colbert_small
scifact_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_answerai_colbert_small.fiqa_answerai_colbert_small
fiqa_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR fiqa (beir/fiqa/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/fiqa/test (ir_datasets 0.6.3), which downloads fiqa.zip (md5 17918ed23cd04fb15047f73e6c3bd9d9). BEIR also publishes this corpus on the Hub as BeIR/fiqa, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from… See the full description on the dataset page: https://huggingface.co/datasets/robro612/fiqa_answerai_colbert_small.nfcorpus_colbertv2
NFCorpus, ColBERTv2
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
Source
BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding
3,633 documents, 323 queries, 12,334 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_colbertv2.nfcorpus_answerai_colbert_small
nfcorpus_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_answerai_colbert_small.scidocs_answerai_colbert_small
scidocs_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scidocs_answerai_colbert_small.lotte_pooled_dev_search_answerai_colbert_small
lotte_pooled_dev_search_answerai_colbert_small
Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_answerai_colbert_small.msmarco_answerai_colbert_small
msmarco_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR msmarco (beir/msmarco/dev), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/msmarco/dev (ir_datasets 0.6.3), which downloads msmarco.zip (md5 444067daf65d982533ea17ebd59501e4). BEIR also publishes this corpus on the Hub as BeIR/msmarco, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/msmarco_answerai_colbert_small.trec-covid_answerai_colbert_small
trec-covid_answerai_colbert_small
Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with
lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.
Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_answerai_colbert_small.ColBERT_Humor_Detection
ColBERT_Humor
Dataset Summary
ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.nfcorpus_modern_colbert
NFCorpus, GTE-ModernColBERT
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
Source
BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding
3,633 documents, 323 queries, 12,334 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_modern_colbert.lotte_pooled_colbertv2
LoTTE pooled (dev), ColBERTv2
Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries.
Source
Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages
Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels)
Document order: passage id order (row i is pid i)
Encoding
Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
Documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/lotte_pooled_colbertv2.ODQA_colbert_top5_100wordsmsmarco-train-distil-colbert-v2miriad-mlateon-colbert-smoke
MIRIAD 200, encoded with mLateOn-medical
Multi-vector (ColBERT-style) embeddings for
tomaarsen/miriad-benchmark-200k,
produced with multi-vector-encoder/mLateOn-medical.
passages
200
token vectors
176,014
mean vectors / passage
880.07
dim
128
stored dtype
float16
embeddings size
0.05 GB
raw text encoded
1 MB
The embeddings are 49x larger than the text
they came from, which is why late-interaction retrieval needs quantization or pooling.… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/miriad-mlateon-colbert-smoke.NQ-colbert-10kopenbookqa_retrieved_by_colbert
Dataset Card for "openbookqa_retrieved_by_colbert"
This is the main/test set of OBQA, with each question retrieved from ColBERT v2 trained on MS MARCO Passage Ranking (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz).
We index the question part of the train set using doc_maxlen=30, nbits=2. We search each question of test set with k=10 and put the results in the retrieved column.
NQ-colbert-10k-casearxiv_colbert_pgtr_golden
ArXiv ColBERT PGTR Golden
ArXiv queries and corpus from Nithish2410/arxiv_colbert_pgtr_golden, with the previous targets ignored and replaced by Qwen-reranked top-100 targets.
Contents
train.jsonl: 10,000 queries with 100 Qwen-reranked targets each.
items.jsonl: 2,040 ArXiv corpus passages.
Rerank Setup
Query source: existing query texts from the dataset.
Corpus source: existing items split from the dataset.
Candidate source: e5-base-v2 retrieval… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/arxiv_colbert_pgtr_golden.commonsense_qa_retrieved_by_colbert
Dataset Card for "commonsense_qa_retrieved_by_colbert"
This is the validation set of CSQA, with each question retrieved from ColBERT v2 trained on MS MARCO Passage Ranking (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz).
We index the question part of the train set using doc_maxlen=30, nbits=2. We search each question of validation set with k=10 and put the results in the retrieved column.
NQ-colbert-20kscidocs_colbert_pgtr_golden
SciDocs ColBERT PGTR Golden
SciDocs queries and corpus from Nithish2410/scidocs_colbert_pgtr_golden, with the previous targets ignored and replaced by full Qwen-reranked top-100 targets.
Contents
train.jsonl: 14,142 queries with 100 Qwen-reranked targets each.
items.jsonl: 25,657 SciDocs corpus passages.
Rerank Setup
Query source: existing query texts from the dataset.
Corpus source: existing items split from the dataset.
Candidate source:… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/scidocs_colbert_pgtr_golden.NQ-colbertcovid_colbert_pgtr_golden
COVID ColBERT PGTR Golden
COVID queries and corpus from Nithish2410/covid_colbert_pgtr_golden, with the previous targets ignored and replaced by Qwen-reranked top-100 targets.
Contents
train.jsonl: 10,000 queries with 100 Qwen-reranked targets each.
items.jsonl: 171,332 COVID corpus passages.
Rerank Setup
Query source: existing query texts from the dataset.
Corpus source: existing items split from the dataset.
Candidate source: e5-base-v2… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/covid_colbert_pgtr_golden.
