Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WenxingZhu /msmarco_answerai_colbert_small_embeddings MS MARCO ColBERT Embeddings Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1. Dataset Structure The dataset contains: data/corpus/: 177 parquet files with document embeddings data/queries/: 11 parquet files with query embeddings data/qrels/train.parquet: Relevance judgments (532,751 pairs) Usage from datasets import load_dataset # Load from directory (recommended for large datasets) corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.tabularfeature-extraction1M<n<10M0 likes213 downloads1y agoHugging Face02johannhartmann /pgturbohybrid_dbpedia_colbert johannhartmann/pgturbohybrid_dbpedia_colbert Precomputed DBpedia ColBERT multivectors for pgturbohybrid benchmark runs. The dataset stores packed little-endian float16 values for the document and query embeddings. Importing these rows into PostgreSQL avoids llama.cpp embedding generation during retrieval/index benchmarks. This compact export stores half-precision values and reconstructs turbohybrid_multivector values on import. Use it for benchmark loading where avoiding runtime… See the full description on the dataset page: https://huggingface.co/datasets/johannhartmann/pgturbohybrid_dbpedia_colbert.tabulartext-retrieval1M<n<10M0 likes206 downloads4mo agoHugging Face03tuskanny /fiqa_modern_colbert FiQA-2018, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding 57,638 documents, 648 queries, 1,706 qrels Text given to the encoder for each document: the passage text (FiQA documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_modern_colbert.tabulartext-retrieval1K<n<10K0 likes145 downloads17d agoHugging Face04tuskanny /scidocs_modern_colbert SCIDOCS, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding 25,657 documents, 1,000 queries, 29,928 qrels Text given to the encoder for each document: title + " " + text (BEIR… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_modern_colbert.tabulartext-retrieval10K<n<100K0 likes136 downloads17d agoHugging Face05tuskanny /scidocs_colbertv2 SCIDOCS, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding 25,657 documents, 1,000 queries, 29,928 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_colbertv2.tabulartext-retrieval10K<n<100K0 likes130 downloads17d agoHugging Face06robro612 /scifact_answerai_colbert_small scifact_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_answerai_colbert_small.tabularn<1K0 likes127 downloads11d agoHugging Face07robro612 /fiqa_answerai_colbert_small fiqa_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR fiqa (beir/fiqa/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/fiqa/test (ir_datasets 0.6.3), which downloads fiqa.zip (md5 17918ed23cd04fb15047f73e6c3bd9d9). BEIR also publishes this corpus on the Hub as BeIR/fiqa, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from… See the full description on the dataset page: https://huggingface.co/datasets/robro612/fiqa_answerai_colbert_small.tabular1K<n<10K0 likes115 downloads11d agoHugging Face08tuskanny /nfcorpus_colbertv2 NFCorpus, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_colbertv2.tabulartext-retrieval10K<n<100K0 likes113 downloads17d agoHugging Face09robro612 /nfcorpus_answerai_colbert_small nfcorpus_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_answerai_colbert_small.tabular10K<n<100K0 likes100 downloads11d agoHugging Face10robro612 /scidocs_answerai_colbert_small scidocs_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scidocs_answerai_colbert_small.tabular10K<n<100K0 likes99 downloads11d agoHugging Face11robro612 /lotte_pooled_dev_search_answerai_colbert_small lotte_pooled_dev_search_answerai_colbert_small Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_answerai_colbert_small.tabular1K<n<10K0 likes98 downloads11d agoHugging Face12robro612 /msmarco_answerai_colbert_small msmarco_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR msmarco (beir/msmarco/dev), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/msmarco/dev (ir_datasets 0.6.3), which downloads msmarco.zip (md5 444067daf65d982533ea17ebd59501e4). BEIR also publishes this corpus on the Hub as BeIR/msmarco, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/msmarco_answerai_colbert_small.tabular1K<n<10K0 likes95 downloads10d agoHugging Face13tuskanny /ms_marco_colbertv2 MS MARCO v1 Passage, ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries. Source Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/ms_marco_colbertv2.tabulartext-retrieval1K<n<10K0 likes88 downloads17d agoHugging Face14robro612 /trec-covid_answerai_colbert_small trec-covid_answerai_colbert_small Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590. Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_answerai_colbert_small.tabular10K<n<100K0 likes87 downloads11d agoHugging Face15tuskanny /fiqa_colbertv2 FiQA-2018, ColBERTv2 Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format. Source BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding 57,638 documents, 648 queries, 1,706 qrels Text given to the encoder for each document: the passage text (FiQA documents have no title). The… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_colbertv2.tabulartext-retrieval1K<n<10K0 likes81 downloads17d agoHugging Face16tuskanny /nfcorpus_modern_colbert NFCorpus, GTE-ModernColBERT Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format. Source BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding 3,633 documents, 323 queries, 12,334 qrels Text given to the encoder for each document: title + " " + text (BEIR title and… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_modern_colbert.tabulartext-retrieval10K<n<100K0 likes77 downloads17d agoHugging Face17tuskanny /lotte_pooled_colbertv2 LoTTE pooled (dev), ColBERTv2 Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries. Source Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels) Document order: passage id order (row i is pid i) Encoding Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer) Documents… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/lotte_pooled_colbertv2.tabulartext-retrieval1K<n<10K0 likes61 downloads17d agoHugging Face18abhinand /clini-colbert-pairs-dev-v2tabular1K<n<10K2 likes8 downloads2y agoHugging Face19Atipico1 /NQ-colbert-10k-case-entitytabular1K<n<10K0 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.