Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Feargal /colbert-retrieval-mined-examplesgated Cantivia internal retrieval training artifacts Proprietary. All rights reserved. This repository stores working artifacts of Cantivia's reranker training pipeline: mined candidate pools, hard negatives, group caches, relevance scores and their manifests. It is published only so Cantivia's own jobs can reach it; it is not a public dataset, and nothing here is released for download, redistribution, derivative works or model training. Licence Cantivia's own… See the full description on the dataset page: https://huggingface.co/datasets/Feargal/colbert-retrieval-mined-examples.0 likes2k downloads2d agoHugging Face02colbertv2 /lotte_passagesLoTTE Passages Dataset for ColBERTv2question-answering1M<n<10M3 likes783 downloads3y agoHugging Face03colbertv2 /lotteLoTTE Passages Dataset for ColBERTv2textquestion-answering10K<n<100K11 likes351 downloads4y agoHugging Face04WenxingZhu /msmarco_answerai_colbert_small_embeddings MS MARCO ColBERT Embeddings Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1. Dataset Structure The dataset contains: data/corpus/: 177 parquet files with document embeddings data/queries/: 11 parquet files with query embeddings data/qrels/train.parquet: Relevance judgments (532,751 pairs) Usage from datasets import load_dataset # Load from directory (recommended for large datasets) corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.tabularfeature-extraction1M<n<10M0 likes301 downloads1y agoHugging Face05jonghwi /msmarco_token_score_colbertx_xlmr_large_zs_en_en Dataset Card for "msmarco_token_score_colbertx_xlmr_large_zs_en_en" More Information needed 1M<n<10M0 likes270 downloads2y agoHugging Face06darvog /hotpotqa_colbert0 likes199 downloads2y agoHugging Face07nielsgl /colbert-wiki20170 likes170 downloads11mo agoHugging Face08darvog /nq_colbert0 likes166 downloads2y agoHugging Face09tuskanny /fiqa_modern_colbert FiQA-2018, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding 57,638 documents, 648 queries, 1,706 qrels Text given to the encoder for each document: the passage text (FiQA documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_modern_colbert.tabulartext-retrieval1K<n<10K0 likes137 downloads13d agoHugging Face10tuskanny /scidocs_modern_colbert SCIDOCS, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding 25,657 documents, 1,000 queries, 29,928 qrels Text given to the encoder for each document: title + " " + text (BEIR… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_modern_colbert.tabulartext-retrieval10K<n<100K0 likes122 downloads13d agoHugging Face11tuskanny /scidocs_colbertv2 SCIDOCS, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding 25,657 documents, 1,000 queries, 29,928 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_colbertv2.tabulartext-retrieval10K<n<100K0 likes114 downloads13d agoHugging Face12robro612 /scifact_answerai_colbert_small scifact_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_answerai_colbert_small.tabularn<1K0 likes108 downloads7d agoHugging Face13yaooooo233 /msmarco-v2-colbertv2-fp32 MS MARCO v2 — ColBERTv2 fp32 Multi-Vector Embeddings (uncompressed) Per-token multi-vector embeddings for the full MS MARCO v2 passage corpus (~138.4M passages) plus the dev / dev2 queries, produced with the official ColBERTv2 checkpoint (colbert-ir/colbertv2.0). Precision: fp32 (NO residual quantization, NO pooling) Dim: 128 per token Corpus token vectors: ~9.41B (avg ~68 tokens/passage) Total corpus size: ~4.82 TB These embeddings are uncompressed on purpose (research on… See the full description on the dataset page: https://huggingface.co/datasets/yaooooo233/msmarco-v2-colbertv2-fp32.text-retrievaln>1T0 likes99 downloads4mo agoHugging Face14karynaur /xpr_colberttext1M<n<10M0 likes97 downloads3y agoHugging Face15eunseong /nq_colbertv20 likes93 downloads1y agoHugging Face16tuskanny /nfcorpus_colbertv2 NFCorpus, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_colbertv2.tabulartext-retrieval10K<n<100K0 likes91 downloads13d agoHugging Face17pyterrier /msmarco_psg_v1.colbertv2 msmarco_psg_v1.colbertv2 Description This is a ColBERT v2 index created via pyterrier_colbert2. This is the index that is used in the SIGIR 2026 PLAID-PRF paper. Usage # Download and load the index from HuggingFace import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/msmarco_psg_v1.colbertv2', plaid_mode=True, ncells=4, centroid_score_threshold=0.4, ndocs=4096) # TREC-DL 2019 pt.Experiment: from pyterrier_colbert.ranking import _prf… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/msmarco_psg_v1.colbertv2.text-retrieval0 likes89 downloads3mo agoHugging Face18CreativeLang /ColBERT_Humor_Detection ColBERT_Humor Dataset Summary ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.text100K<n<1M7 likes86 downloads3y agoHugging Face19tuskanny /ms_marco_colbertv2 MS MARCO v1 Passage, ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries. Source Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/ms_marco_colbertv2.tabulartext-retrieval1K<n<10K0 likes84 downloads13d agoHugging Face20robro612 /fiqa_answerai_colbert_small fiqa_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR fiqa (beir/fiqa/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/fiqa/test (ir_datasets 0.6.3), which downloads fiqa.zip (md5 17918ed23cd04fb15047f73e6c3bd9d9). BEIR also publishes this corpus on the Hub as BeIR/fiqa, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from… See the full description on the dataset page: https://huggingface.co/datasets/robro612/fiqa_answerai_colbert_small.tabular1K<n<10K0 likes82 downloads7d agoHugging Face21robro612 /trec-covid_answerai_colbert_small trec-covid_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_answerai_colbert_small.tabular10K<n<100K0 likes79 downloads7d agoHugging Face22robro612 /lotte_pooled_dev_search_answerai_colbert_small lotte_pooled_dev_search_answerai_colbert_small Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_answerai_colbert_small.tabular1K<n<10K0 likes78 downloads7d agoHugging Face23robro612 /nfcorpus_answerai_colbert_small nfcorpus_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_answerai_colbert_small.tabular10K<n<100K0 likes75 downloads7d agoHugging Face24robro612 /scidocs_answerai_colbert_small scidocs_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scidocs_answerai_colbert_small.tabular10K<n<100K0 likes73 downloads7d agoHugging Face25robro612 /msmarco_answerai_colbert_small msmarco_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR msmarco (beir/msmarco/dev), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/msmarco/dev (ir_datasets 0.6.3), which downloads msmarco.zip (md5 444067daf65d982533ea17ebd59501e4). BEIR also publishes this corpus on the Hub as BeIR/msmarco, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/msmarco_answerai_colbert_small.tabular1K<n<10K0 likes73 downloads7d agoHugging Face26tuskanny /fiqa_colbertv2 FiQA-2018, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding 57,638 documents, 648 queries, 1,706 qrels Text given to the encoder for each document: the passage text (FiQA documents have no title). The… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_colbertv2.tabulartext-retrieval1K<n<10K0 likes72 downloads13d agoHugging Face27tuskanny /lotte_pooled_colbertv2 LoTTE pooled (dev), ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries. Source Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer) Documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/lotte_pooled_colbertv2.tabulartext-retrieval1K<n<10K0 likes65 downloads13d agoHugging Face28tuskanny /nfcorpus_modern_colbert NFCorpus, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_modern_colbert.tabulartext-retrieval10K<n<100K0 likes59 downloads13d agoHugging Face29jpq-repro /msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__ msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__ Description This is the PyTerrier JPQIndex for MSMARCO v1 passage corpus, which corresponds to an result from the SIGIR 2026 reproducibility paper. Usage # Load the artifact import pyterrier as pt import pyterrier_dr index = pt.Artifact.from_hf('jpq-repro/msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__') model =… See the full description on the dataset page: https://huggingface.co/datasets/jpq-repro/msmarco-passage-train__tct_colbert__faiss2opq__M96_nbits8__ps159744__neg200__ibn__lr__.text-retrieval0 likes45 downloads3mo agoHugging Face30yosefw /msmarco-train-distil-colbert-v2text100K<n<1M0 likes44 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.