Team Ai
20 results

ColBERT

Feargal /colbert-retrieval-mined-examplesgated Cantivia internal retrieval training artifacts Proprietary. All rights reserved. This repository stores working artifacts of Cantivia's reranker training pipeline: mined candidate pools, hard negatives, group caches, relevance scores and their manifests. It is published only so Cantivia's own jobs can reach it; it is not a public dataset, and nothing here is released for download, redistribution, derivative works or model training. Licence Cantivia's own… See the full description on the dataset page: https://huggingface.co/datasets/Feargal/colbert-retrieval-mined-examples.0 likes2k downloads1d agoHugging Facecolbertv2 /lotte_passagesLoTTE Passages Dataset for ColBERTv2question-answering1M<n<10M3 likes761 downloads3y agoHugging FaceWenxingZhu /msmarco_answerai_colbert_small_embeddings MS MARCO ColBERT Embeddings Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1. Dataset Structure The dataset contains: data/corpus/: 177 parquet files with document embeddings data/queries/: 11 parquet files with query embeddings data/qrels/train.parquet: Relevance judgments (532,751 pairs) Usage from datasets import load_dataset # Load from directory (recommended for large datasets) corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.tabularfeature-extraction1M<n<10M0 likes413 downloads1y agoHugging Facecolbertv2 /lotteLoTTE Passages Dataset for ColBERTv2textquestion-answering10K<n<100K11 likes344 downloads4y agoHugging Facejonghwi /msmarco_token_score_colbertx_xlmr_large_zs_en_en Dataset Card for "msmarco_token_score_colbertx_xlmr_large_zs_en_en" More Information needed 1M<n<10M0 likes279 downloads2y agoHugging Facedarvog /hotpotqa_colbert0 likes199 downloads2y agoHugging Face