Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01colbertv2 /lotteLoTTE Passages Dataset for ColBERTv2textquestion-answering10K<n<100K11 likes303 downloads4y agoHugging Face02WenxingZhu /msmarco_answerai_colbert_small_embeddings MS MARCO ColBERT Embeddings Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1. Dataset Structure The dataset contains: data/corpus/: 177 parquet files with document embeddings data/queries/: 11 parquet files with query embeddings data/qrels/train.parquet: Relevance judgments (532,751 pairs) Usage from datasets import load_dataset # Load from directory (recommended for large datasets) corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.tabularfeature-extraction1M<n<10M0 likes213 downloads1y agoHugging Face03johannhartmann /pgturbohybrid_dbpedia_colbert johannhartmann/pgturbohybrid_dbpedia_colbert Precomputed DBpedia ColBERT multivectors for pgturbohybrid benchmark runs. The dataset stores packed little-endian float16 values for the document and query embeddings. Importing these rows into PostgreSQL avoids llama.cpp embedding generation during retrieval/index benchmarks. This compact export stores half-precision values and reconstructs turbohybrid_multivector values on import. Use it for benchmark loading where avoiding runtime… See the full description on the dataset page: https://huggingface.co/datasets/johannhartmann/pgturbohybrid_dbpedia_colbert.tabulartext-retrieval1M<n<10M0 likes206 downloads4mo agoHugging Face04karynaur /xpr_colberttext1M<n<10M0 likes178 downloads3y agoHugging Face05KShivendu /nanobeir-colbert-token-vectors NanoBEIR ColBERT token vectors Per-token ColBERT vectors for all 13 NanoBEIR datasets, from two late-interaction models: folder model revision encoder dtype lateon/ lightonai/LateOn 62911e105059585d244384c7d17826e35f669c17 fp32 iso/ topk-io/Iso-ModernColBERT b95d9608ad424fdd7bfd045001576f59cbb89a98 bf16 pplx-late-0.6b/ perplexity-ai/pplx-embed-v2-late-0.6b dd4e95b836a73f6f0c32e46ea127c0b86b02169e bf16 Each model covers 56,723 documents (7,585,402 kept… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/nanobeir-colbert-token-vectors.texttext-retrieval100K<n<1M0 likes161 downloads2d agoHugging Face06tuskanny /scidocs_modern_colbert SCIDOCS, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding 25,657 documents, 1,000 queries, 29,928 qrels Text given to the encoder for each document: title + " " + text (BEIR… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_modern_colbert.tabulartext-retrieval10K<n<100K0 likes136 downloads17d agoHugging Face07tuskanny /scidocs_colbertv2 SCIDOCS, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding 25,657 documents, 1,000 queries, 29,928 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_colbertv2.tabulartext-retrieval10K<n<100K0 likes130 downloads17d agoHugging Face08robro612 /scifact_answerai_colbert_small scifact_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_answerai_colbert_small.tabularn<1K0 likes127 downloads11d agoHugging Face09robro612 /fiqa_answerai_colbert_small fiqa_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR fiqa (beir/fiqa/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/fiqa/test (ir_datasets 0.6.3), which downloads fiqa.zip (md5 17918ed23cd04fb15047f73e6c3bd9d9). BEIR also publishes this corpus on the Hub as BeIR/fiqa, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from… See the full description on the dataset page: https://huggingface.co/datasets/robro612/fiqa_answerai_colbert_small.tabular1K<n<10K0 likes115 downloads11d agoHugging Face10tuskanny /nfcorpus_colbertv2 NFCorpus, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_colbertv2.tabulartext-retrieval10K<n<100K0 likes113 downloads17d agoHugging Face11robro612 /nfcorpus_answerai_colbert_small nfcorpus_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_answerai_colbert_small.tabular10K<n<100K0 likes100 downloads11d agoHugging Face12robro612 /scidocs_answerai_colbert_small scidocs_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scidocs_answerai_colbert_small.tabular10K<n<100K0 likes99 downloads11d agoHugging Face13robro612 /lotte_pooled_dev_search_answerai_colbert_small lotte_pooled_dev_search_answerai_colbert_small Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_answerai_colbert_small.tabular1K<n<10K0 likes98 downloads11d agoHugging Face14robro612 /msmarco_answerai_colbert_small msmarco_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR msmarco (beir/msmarco/dev), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/msmarco/dev (ir_datasets 0.6.3), which downloads msmarco.zip (md5 444067daf65d982533ea17ebd59501e4). BEIR also publishes this corpus on the Hub as BeIR/msmarco, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/msmarco_answerai_colbert_small.tabular1K<n<10K0 likes95 downloads10d agoHugging Face15robro612 /trec-covid_answerai_colbert_small trec-covid_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_answerai_colbert_small.tabular10K<n<100K0 likes87 downloads11d agoHugging Face16CreativeLang /ColBERT_Humor_Detection ColBERT_Humor Dataset Summary ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.text100K<n<1M7 likes79 downloads3y agoHugging Face17tuskanny /nfcorpus_modern_colbert NFCorpus, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_modern_colbert.tabulartext-retrieval10K<n<100K0 likes77 downloads17d agoHugging Face18tuskanny /lotte_pooled_colbertv2 LoTTE pooled (dev), ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries. Source Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer) Documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/lotte_pooled_colbertv2.tabulartext-retrieval1K<n<10K0 likes61 downloads17d agoHugging Face19seonjeongh /ODQA_colbert_top5_100wordstext10K<n<100K0 likes36 downloads2y agoHugging Face20yosefw /msmarco-train-distil-colbert-v2text100K<n<1M0 likes35 downloads1y agoHugging Face21KShivendu /miriad-mlateon-colbert-smoke MIRIAD 200, encoded with mLateOn-medical Multi-vector (ColBERT-style) embeddings for tomaarsen/miriad-benchmark-200k, produced with multi-vector-encoder/mLateOn-medical. passages 200 token vectors 176,014 mean vectors / passage 880.07 dim 128 stored dtype float16 embeddings size 0.05 GB raw text encoded 1 MB The embeddings are 49x larger than the text they came from, which is why late-interaction retrieval needs quantization or pooling.… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/miriad-mlateon-colbert-smoke.texttext-retrievaln<1K0 likes32 downloads1mo agoHugging Face22Atipico1 /NQ-colbert-10ktext10K<n<100K0 likes30 downloads3y agoHugging Face23cnut1648 /openbookqa_retrieved_by_colbert Dataset Card for "openbookqa_retrieved_by_colbert" This is the main/test set of OBQA, with each question retrieved from ColBERT v2 trained on MS MARCO Passage Ranking (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz). We index the question part of the train set using doc_maxlen=30, nbits=2. We search each question of test set with k=10 and put the results in the retrieved column. textn<1K1 likes14 downloads3y agoHugging Face24Atipico1 /NQ-colbert-10k-casetext10K<n<100K0 likes13 downloads3y agoHugging Face25Nithish2410 /arxiv_colbert_pgtr_golden ArXiv ColBERT PGTR Golden ArXiv queries and corpus from Nithish2410/arxiv_colbert_pgtr_golden, with the previous targets ignored and replaced by Qwen-reranked top-100 targets. Contents train.jsonl: 10,000 queries with 100 Qwen-reranked targets each. items.jsonl: 2,040 ArXiv corpus passages. Rerank Setup Query source: existing query texts from the dataset. Corpus source: existing items split from the dataset. Candidate source: e5-base-v2 retrieval… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/arxiv_colbert_pgtr_golden.text10K<n<100K0 likes13 downloads1mo agoHugging Face26cnut1648 /commonsense_qa_retrieved_by_colbert Dataset Card for "commonsense_qa_retrieved_by_colbert" This is the validation set of CSQA, with each question retrieved from ColBERT v2 trained on MS MARCO Passage Ranking (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz). We index the question part of the train set using doc_maxlen=30, nbits=2. We search each question of validation set with k=10 and put the results in the retrieved column. text1K<n<10K0 likes11 downloads3y agoHugging Face27Atipico1 /NQ-colbert-20ktext10K<n<100K0 likes11 downloads3y agoHugging Face28Nithish2410 /scidocs_colbert_pgtr_golden SciDocs ColBERT PGTR Golden SciDocs queries and corpus from Nithish2410/scidocs_colbert_pgtr_golden, with the previous targets ignored and replaced by full Qwen-reranked top-100 targets. Contents train.jsonl: 14,142 queries with 100 Qwen-reranked targets each. items.jsonl: 25,657 SciDocs corpus passages. Rerank Setup Query source: existing query texts from the dataset. Corpus source: existing items split from the dataset. Candidate source:… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/scidocs_colbert_pgtr_golden.text10K<n<100K0 likes11 downloads2mo agoHugging Face29Atipico1 /NQ-colberttext10K<n<100K0 likes10 downloads3y agoHugging Face30Nithish2410 /covid_colbert_pgtr_golden COVID ColBERT PGTR Golden COVID queries and corpus from Nithish2410/covid_colbert_pgtr_golden, with the previous targets ignored and replaced by Qwen-reranked top-100 targets. Contents train.jsonl: 10,000 queries with 100 Qwen-reranked targets each. items.jsonl: 171,332 COVID corpus passages. Rerank Setup Query source: existing query texts from the dataset. Corpus source: existing items split from the dataset. Candidate source: e5-base-v2… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/covid_colbert_pgtr_golden.text100K<n<1M0 likes9 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.