Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fineweb-retrieval /fineweb-edu-indexThis dataset contains the embeddings for the full fineweb-edu, embedded with the Cohere Embed V3 model. You can search on this dataset with just 500MB of memory using DiskVectorIndex. Installation & Usage Get your free Cohere API key from cohere.com. You must set this API key as an environment variable: export COHERE_API_KEY=your_api_key Install the package: pip install DiskVectorIndex You can then search via: from DiskVectorIndex import DiskVectorIndex index =… See the full description on the dataset page: https://huggingface.co/datasets/fineweb-retrieval/fineweb-edu-index.0 likes23k downloads1y agoHugging Face02CoIR-Retrieval /cosqaEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/cosqa.text10K<n<100K0 likes4.9k downloads2y agoHugging Face03CoIR-Retrieval /appsEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/apps.text10K<n<100K1 likes4.6k downloads2y agoHugging Face04BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes4.2k downloads1y agoHugging Face05nlphuji /flickr_1k_test_image_text_retrieval Flickr30k (1K test set) Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions Homepage: https://shannon.cs.illinois.edu/DenotationGraph/ 1K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip Bibtex: @article{young2014image, title={From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/flickr_1k_test_image_text_retrieval.image1K<n<10K2 likes3.1k downloads4y agoHugging Face06PaDaS-Lab /webfaq-retrievalWebFAQ Retrieval Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages. Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.texttext-retrieval10M<n<100M10 likes3k downloads1y agoHugging Face07CoIR-Retrieval /stackoverflow-qaEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/stackoverflow-qa.text10K<n<100K0 likes2.8k downloads2y agoHugging Face08CoIR-Retrieval /CodeSearchNetEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/CodeSearchNet.text1M<n<10M3 likes2.7k downloads2y agoHugging Face09nlphuji /mscoco_2014_5k_test_image_text_retrieval MSCOCO (5K test set) Original paper: Microsoft COCO: Common Objects in Context Homepage: https://cocodataset.org/#home 5K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip Bibtex: @inproceedings{lin2014microsoft, title={Microsoft coco: Common objects in context}, author={Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Doll{\'a}r, Piotr and Zitnick, C Lawrence}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/mscoco_2014_5k_test_image_text_retrieval.image1K<n<10K11 likes2.2k downloads4y agoHugging Face10CoIR-Retrieval /CodeSearchNet-ccrEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/CodeSearchNet-ccr.text1M<n<10M1 likes2k downloads2y agoHugging Face11Feargal /colbert-retrieval-mined-examplesgated Cantivia internal retrieval training artifacts Proprietary. All rights reserved. This repository stores working artifacts of Cantivia's reranker training pipeline: mined candidate pools, hard negatives, group caches, relevance scores and their manifests. It is published only so Cantivia's own jobs can reach it; it is not a public dataset, and nothing here is released for download, redistribution, derivative works or model training. Licence Cantivia's own… See the full description on the dataset page: https://huggingface.co/datasets/Feargal/colbert-retrieval-mined-examples.0 likes1.9k downloads4h agoHugging Face12TableQAKit /FinQA_retrieval0 likes1.9k downloads3y agoHugging Face13CoIR-Retrieval /codefeedback-stEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/codefeedback-st.text100K<n<1M0 likes1.9k downloads2y agoHugging Face14CoIR-Retrieval /synthetic-text2sqlEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/synthetic-text2sql.text100K<n<1M0 likes1.4k downloads2y agoHugging Face15CoIR-Retrieval /codetrans-dlEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/codetrans-dl.text1K<n<10K0 likes1.4k downloads2y agoHugging Face16vnahata /IndicDiarBench-speaker-retrieval Indic DiarBench speaker retrieval (MTEB) Indic DiarBench reshaped for speaker retrieval across the 22 scheduled languages of India: given a clip of one speaker, find other clips of that same speaker. Source: sarvamai/indic-diarbench at revision 92877ba, cc-by-4.0, official test split. Turns are cut by their annotated times, restricted to 2 to 15 seconds, and turns overlapping a different speaker are dropped. identity pairs the recording session with the speaker, because speaker… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/IndicDiarBench-speaker-retrieval.audioaudio-classification1K<n<10K0 likes1.3k downloads1mo agoHugging Face17CoIR-Retrieval /codefeedback-mtEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/codefeedback-mt.text100K<n<1M0 likes1.2k downloads2y agoHugging Face18CoIR-Retrieval /codetrans-contestEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/codetrans-contest.text1K<n<10K0 likes1.2k downloads2y agoHugging Face19vnahata /OmnilingualASR-retrieval Omnilingual ASR speech-text retrieval (MTEB) Read speech paired with its human transcription, for languages that no existing MTEB audio task covers. Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated transcripts are dropped, since one would otherwise be relevant to several recordings while only one is marked correct. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.audioautomatic-speech-recognition1K<n<10K0 likes1.2k downloads1mo agoHugging Face20irodkin /kv_retrievaltext10M<n<100M0 likes1.1k downloads2mo agoHugging Face21TableQAKit /MultiHiertt_retrieval0 likes1.1k downloads3y agoHugging Face22vnahata /AfriMCQA-qa-retrieval Afri-MCQA visual QA retrieval (MTEB) Culturally grounded multiple choice questions about photographs, in 16 African languages. A query is a question with its photograph; the corpus holds every distinct answer option in that language, so a question is ranked against the whole option pool rather than only its own four. Built from Atnafu/Afri-MCQA at revision 8b8c53d, cc-by-nc-4.0, using the official dev split, the only one where the correct answer is labelled. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-qa-retrieval.imagevisual-question-answering10K<n<100K0 likes1k downloads1mo agoHugging Face23jinaai /github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.image1K<n<10K0 likes956 downloads1y agoHugging Face24kasys /multimodal-tip-of-the-tongue-retrieval-for-scientific-documents Open-Source Scientific Documents This repository contains the scientific-paper corpus and generated artifacts for Multimodal Tip-of-the-Tongue Retrieval for Scientific Papers. It brings together source PDFs, extracted paper content, textual and visual clues, query collections, and split definitions. The generation code documents how these artifacts are made. The PDFs form a shared retrieval corpus. Query and evaluation collections refer to paper identifiers in that corpus; the… See the full description on the dataset page: https://huggingface.co/datasets/kasys/multimodal-tip-of-the-tongue-retrieval-for-scientific-documents.documentvisual-document-retrieval100K<n<1M1 likes874 downloads4d agoHugging Face25sentence-transformers /quantized-retrieval-datatext10M<n<100M2 likes783 downloads9mo agoHugging Face26ai-forever /rubq-retrievaltexttext-retrieval10K<n<100K2 likes777 downloads2y agoHugging Face27MatanBT /retrieval-datasets-similarities Summary Caching the similarity results of different embedding-based retrieval, on different dataset; that is, the similarities between each query and all the corpus passages. Method. We collect these results in JSON files, containing the similarities similarities that were collected running evaluation with (BEIR), on the specific model and data. Full list below. Usage. This caching can be used to evaluate the benign accuracy of the models, and---more importantly---to explore the… See the full description on the dataset page: https://huggingface.co/datasets/MatanBT/retrieval-datasets-similarities.0 likes748 downloads1y agoHugging Face28jinaai /airbnb-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval_beir.image1K<n<10K0 likes739 downloads1y agoHugging Face29fineweb-retrieval /fineweb-edu-embThis file contains the embeddings for the full fineweb-edu dataset. The dataset has been deduplicated (using only exact deduplication). The emb folder contains for each parquet file a new_{parquet_name}.npy and old_{parquet}.npy file. The old refers to text that has been seen in the smaller 10B/100B/350B data samples. Cohere embed-multilingual-v3.0 model has been used. The index of the dataset can be found here: https://huggingface.co/datasets/Cohere/fineweb-edu-index The corpus can be found… See the full description on the dataset page: https://huggingface.co/datasets/fineweb-retrieval/fineweb-edu-emb.4 likes716 downloads2y agoHugging Face30Snowflake /mteb-retrieval-snowflake-arctic-embed-m-v1.5text10M<n<100M0 likes699 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.