Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BeIR /scifact Dataset Card for BEIR Benchmark scifact is one of the datasets from the Fact Checking task within BEIR, measuring scientific article retrieval for a given scientific claim. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact.textzero-shot-classification1K<n<10K9 likes25k downloads6mo agoHugging Face02BeIR /scifact-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact-qrels.tabulartext-retrieval1K<n<10K1 likes18k downloads4y agoHugging Face03BeIR /scidocs Dataset Card for BEIR Benchmark scidocs is one of the datasets from the Citation Prediction task within BEIR, measuring cited scientific article retrieval for a given scientific title. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scidocs.textzero-shot-classification10K<n<100K10 likes12k downloads6mo agoHugging Face04BeIR /nfcorpus Dataset Card for BEIR Benchmark nfcorpus is one of the datasets from the Bio-Medical Retrieval task within BEIR, measuring the retrieval of scientific articles for a given query about a nutritional fact. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nfcorpus.textzero-shot-classification1K<n<10K7 likes8k downloads6mo agoHugging Face05pinecone /msmarco-beir-e50 likes8k downloads1y agoHugging Face06BeIR /nfcorpus-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nfcorpus-qrels.texttext-retrieval100K<n<1M0 likes6.8k downloads4y agoHugging Face07castorini /prebuilt-indexes-beir Prebuilt Indexes for BEIR Available indexes: Lucene Flat beir-v1.0.0-trec-covid.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'trec-covid' encoded by BGE-base-en-v1.5. beir-v1.0.0-bioasq.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'bioasq' encoded by BGE-base-en-v1.5. beir-v1.0.0-nfcorpus.bge-base-en-v1.5.flat [readme] Lucene flat index of BEIR collection 'nfcorpus' encoded by BGE-base-en-v1.5. beir-v1.0.0-nq.bge-base-en-v1.5.flat… See the full description on the dataset page: https://huggingface.co/datasets/castorini/prebuilt-indexes-beir.1 likes6.3k downloads1y agoHugging Face08clips /beir-nl-cqadupstack Dataset Card for BEIR-NL Benchmark Dataset Summary BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). BEIR-NL contains the following tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.texttext-retrieval100K<n<1M0 likes6.2k downloads2y agoHugging Face09BeIR /fiqa Dataset Card for BEIR Benchmark fiqa is one of the datasets from the Question Answering task within BEIR, measuring financial article retrieval for a given financial query. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/fiqa.textzero-shot-classification10K<n<100K16 likes4.2k downloads6mo agoHugging Face10BeIR /fiqa-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/fiqa-qrels.tabulartext-retrieval10K<n<100K0 likes3.7k downloads4y agoHugging Face11TIGER-Lab /M-BEIR UniIR: Training and Benchmarking Universal Multimodal Information Retrievers (ECCV 2024) 🌐 Homepage | 🤗 Model(UniIR Checkpoints) | 🤗 Paper | 📖 arXiv | GitHub How to download the M-BEIR Dataset 🔔News 🔥[2023-12-21]: Our M-BEIR Benchmark is now available for use. Dataset Summary M-BEIR, the Multimodal BEnchmark for Instructed Retrieval, is a comprehensive large-scale retrieval benchmark designed to train and evaluate unified multimodal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/M-BEIR.texttext-retrieval1M<n<10M27 likes2.5k downloads2y agoHugging Face12pinecone /msmarco-beir-constbert0 likes2.4k downloads1y agoHugging Face13BeIR /trec-covid Dataset Card for BEIR Benchmark trec-covid is one of the datasets from the Bio-Medical Retrieval task within BEIR, measuring scientific article retrieval for a given query on COVID-19. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid.textzero-shot-classification100K<n<1M7 likes2.4k downloads6mo agoHugging Face14BeIR /quora Dataset Card for BEIR Benchmark quora is one of the datasets from the Duplicate Question Retrieval task within BEIR, measuring duplicate query retrieval for a given query. NOTE: ArguAna has queries also incorporated within the corpus, so you should remove the same query_id if present within the corpus during inference (implemented in BEIR) Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/quora.textzero-shot-classification100K<n<1M5 likes2.1k downloads6mo agoHugging Face15vidore /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K1 likes2.1k downloads1y agoHugging Face16BeIR /msmarco Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. This msmarco subset is part of BEIR. Languages All tasks are in English (en). Dataset Structure This dataset uses the standard BEIR retrieval layout and includes: corpus: one row per document with _id, title, text queries: one row per query with _id, title, text Data Fields _id… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/msmarco.textzero-shot-classification1M<n<10M14 likes2k downloads6mo agoHugging Face17vidore /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes2k downloads1y agoHugging Face18vidore /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes2k downloads1y agoHugging Face19vidore /shiftproject_test_beirBEIR version of vidore/shiftproject_test. imagedocument-question-answering1K<n<10K0 likes1.9k downloads1y agoHugging Face20vidore /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.9k downloads1y agoHugging Face21BeIR /trec-covid-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-qrels.tabulartext-retrieval10K<n<100K1 likes1.9k downloads4y agoHugging Face22vidore /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes1.9k downloads1y agoHugging Face23vidore /tatdqa_test_beirBEIR version of vidore/tatdqa_test. imagedocument-question-answering1K<n<10K0 likes1.9k downloads1y agoHugging Face24vidore /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face25vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face26vidore /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K1 likes1.8k downloads1y agoHugging Face27BeIR /arguana Dataset Card for BEIR Benchmark arguana is one of the datasets from the Argument Retrieval task within BEIR, evaluating retrieving counterarguments for a given argument as query. NOTE: ArguAna has queries also incorporated within the corpus, so you should remove the same query_id if present within corpus during inference (implemented in BEIR) Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/arguana.textzero-shot-classification10K<n<100K4 likes1.8k downloads6mo agoHugging Face28CohereLabs /beir-embed-english-v3 BEIR embeddings with Cohere embed-english-v3.0 model This datasets contains all query & document embeddings for BEIR, embedded with the Cohere embed-english-v3.0 embedding model. Overview of datasets This repository hosts all 18 datasets from BEIR, including query and document embeddings. The following table gives an overview of the available datasets. See the next section how to load the individual datasets. Dataset nDCG@10 #Documents arguana 53.98 8,674… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/beir-embed-english-v3.text10M<n<100M9 likes1.7k downloads7mo agoHugging Face29BeIR /hotpotqa Dataset Card for BEIR Benchmark hotpotqa is one of the datasets from the Question Answering task within BEIR, measuring Wikipedia article retrieval for a given multi-hop query. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/hotpotqa.textzero-shot-classification1M<n<10M17 likes1.6k downloads6mo agoHugging Face30BeIR /nq Dataset Card for BEIR Benchmark nq is one of the datasets from the Question Answering task within BEIR, measuring Wikipedia article retrieval for a web search query. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nq.textzero-shot-classification1M<n<10M4 likes1.5k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.