Team Ai
20 results

colbert

Feargal /colbert-retrieval-mined-examplesgated Cantivia internal retrieval training artifacts Proprietary. All rights reserved. This repository stores working artifacts of Cantivia's reranker training pipeline: mined candidate pools, hard negatives, group caches, relevance scores and their manifests. It is published only so Cantivia's own jobs can reach it; it is not a public dataset, and nothing here is released for download, redistribution, derivative works or model training. Licence Cantivia's own… See the full description on the dataset page: https://huggingface.co/datasets/Feargal/colbert-retrieval-mined-examples.0 likes1.9k downloads8h agoHugging Facecolbertv2 /lotte_passagesLoTTE Passages Dataset for ColBERTv2question-answering1M<n<10M3 likes780 downloads3y agoHugging Facecolbertv2 /lotteLoTTE Passages Dataset for ColBERTv2textquestion-answering10K<n<100K11 likes303 downloads4y agoHugging FaceWenxingZhu /msmarco_answerai_colbert_small_embeddings MS MARCO ColBERT Embeddings Pre-computed ColBERT embeddings for MS MARCO using PyLate and answerdotai/answerai-colbert-small-v1. Dataset Structure The dataset contains: data/corpus/: 177 parquet files with document embeddings data/queries/: 11 parquet files with query embeddings data/qrels/train.parquet: Relevance judgments (532,751 pairs) Usage from datasets import load_dataset # Load from directory (recommended for large datasets) corpus =… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/msmarco_answerai_colbert_small_embeddings.tabularfeature-extraction1M<n<10M0 likes213 downloads1y agoHugging Facejohannhartmann /pgturbohybrid_dbpedia_colbert johannhartmann/pgturbohybrid_dbpedia_colbert Precomputed DBpedia ColBERT multivectors for pgturbohybrid benchmark runs. The dataset stores packed little-endian float16 values for the document and query embeddings. Importing these rows into PostgreSQL avoids llama.cpp embedding generation during retrieval/index benchmarks. This compact export stores half-precision values and reconstructs turbohybrid_multivector values on import. Use it for benchmark loading where avoiding runtime… See the full description on the dataset page: https://huggingface.co/datasets/johannhartmann/pgturbohybrid_dbpedia_colbert.tabulartext-retrieval1M<n<10M0 likes206 downloads4mo agoHugging Facedarvog /hotpotqa_colbert0 likes199 downloads2y agoHugging Face