Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes4.2k downloads1y agoHugging Face02jinaai /github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.image1K<n<10K0 likes956 downloads1y agoHugging Face03jinaai /airbnb-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval_beir.image1K<n<10K0 likes739 downloads1y agoHugging Face04allenai /pointer-retrievaltabular100K<n<1M0 likes696 downloads8mo agoHugging Face05jinaai /tweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.image1K<n<10K0 likes556 downloads1y agoHugging Face06BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes477 downloads2y agoHugging Face07mjbommar /opengloss-v2.4-retrieval-pairs OpenGloss v2.4 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an example paired with its own sense's gloss (positive), and optional sampled cross-headword same-domain negatives. Every pair carries both spans, both reading levels, and live_senses, so a consumer can filter or reweight by… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.4-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes476 downloads16d agoHugging Face08closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes415 downloads4y agoHugging Face09nmixx-fin /twice_tat_qa_retrieval TATQA-Retrieval-ko Constructed a TAT (Textual and Tabular) QA dataset based on information collected from Korean financial reports. tabular1K<n<10K0 likes324 downloads1y agoHugging Face10nmixx-fin /twice_kr_market_report_retrieval FinMarketReport-Retrieval-ko Constructed a Retrieval dataset related to the stock market, based on Korean Financial Reports. tabular1K<n<10K0 likes292 downloads1y agoHugging Face11nmixx-fin /twice_kr_news_retrieval FinNews-Retrieval-ko Constructed a Retrieval dataset based on Korean financial news articles. tabular1K<n<10K0 likes287 downloads1y agoHugging Face12airefinery /esg_cid_retrieval Enhancing Retrieval for ESGLLM via ESG-CID -- A Disclosure Content Index Finetuning Dataset for Mapping GRI and ESRS Usage from datasets import load_dataset # document chunks: train/dev/test_gri/test_esrs documents = load_dataset("esgllm/esg_cid_retrieval", "documents") # queries (disclosure text): train/dev/test_gri/test_esrs queries = load_dataset("esgllm/esg_cid_retrieval", "queries") # training triplets: train/dev triplets = load_dataset("esgllm/esg_cid_retrieval"… See the full description on the dataset page: https://huggingface.co/datasets/airefinery/esg_cid_retrieval.tabular10K<n<100K3 likes259 downloads1y agoHugging Face13mjbommar /opengloss-v2.3-retrieval-pairs Superseded by OpenGloss v2.4 (2026-09-25): every sense now has search queries, QA pairs and verified examples (v2.3 had them only for core and tier 2); level x register definitions and leveled contrasts and explanations are added; and the pretraining corpus no longer contains duplicate documents. v2.3 stays published for reproducibility. OpenGloss v2.3 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes222 downloads16d agoHugging Face14mjbommar /opengloss-v2.1-retrieval-pairs Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes217 downloads1mo agoHugging Face15lighteval /treb_table_retrievaltabularn<1K0 likes188 downloads1y agoHugging Face16closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15image10M<n<100M0 likes169 downloads4y agoHugging Face17Xense /newbalance_shoe_insole_retrieval_and_packing_0611This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_flexiv_rizon4_rt", "total_episodes": 298, "total_frames": 4967405, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:298" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0611.tabularrobotics1M<n<10M0 likes167 downloads3mo agoHugging Face18Self-GRIT /open-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokenstabular1M<n<10M0 likes157 downloads2y agoHugging Face19mjbommar /opengloss-v2.2-retrieval-pairs Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility. OpenGloss v2.2 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes136 downloads1mo agoHugging Face20Xense /shoe_insole_retrieval_and_packing0515This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_flexiv_rizon4_rt", "total_episodes": 101, "total_frames": 206299, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:101" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/shoe_insole_retrieval_and_packing0515.tabularrobotics100K<n<1M0 likes116 downloads5mo agoHugging Face21Xense /newbalance_shoe_insole_retrieval_and_packing_0604This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_flexiv_rizon4_rt", "total_episodes": 108, "total_frames": 1635459, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:108" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0604.tabularrobotics1M<n<10M0 likes111 downloads4mo agoHugging Face22closji /mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15tabular10M<n<100M0 likes110 downloads4y agoHugging Face23Maho001 /tda-gnn-scientific-retrievaltabular1M<n<10M0 likes108 downloads3h agoHugging Face24HZ-VUW /Pathology-Retrieval-Benchmark Pathology Retrieval Benchmark — Rerank Cache Privacy-safe embedding cache for rerank-only reproduction (no document IDs or file paths). Files One Parquet file per query: entry_{id}.parquet (100 entries). Schema Column Description kind query_text, query_image, or candidate candidate_idx 0–19 for candidates; null for queries score Text-retrieval score (candidates only) embedding ColQwen multi-vector embedding (N, 128) Each file… See the full description on the dataset page: https://huggingface.co/datasets/HZ-VUW/Pathology-Retrieval-Benchmark.tabular1K<n<10K0 likes102 downloads4mo agoHugging Face25mjbommar /opengloss-v2.0-retrieval-pairs Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-retrieval-pairs.tabularsentence-similarity1M<n<10M0 likes100 downloads1mo agoHugging Face26closji /mscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01tabular10M<n<100M1 likes97 downloads4y agoHugging Face27nbalepur /karl-with-retrieval_bm25_v2 Dataset Card for "karl-with-retrieval_bm25_v2" More Information needed tabular100K<n<1M0 likes97 downloads2y agoHugging Face28CMiller /kbmill-brick-retrieval KBMill Brick Retrieval Demos Portable, residual-honest knowledge bricks turned into retrieval evaluation sets for KBMill — the public mill at kbmill.com. Shelf packages live in kbmill-brick-library. These are not unbounded wiki dumps or synthetic QA. They come from real KBMill manufacturing: bounded packages with muted residual junk filtered where applicable, craft notes, security report in the ZIP, and published, re-runnable cosine retrieval evidence. Config Queries… See the full description on the dataset page: https://huggingface.co/datasets/CMiller/kbmill-brick-retrieval.tabulartext-retrieval1K<n<10K0 likes93 downloads1mo agoHugging Face29HemanthKari /rcsb-ligand-retrieval-v0.1 The RCSB Protein-Ligand Retrieval Index An openly licensed retrieval dataset of 72,199 protein-ligand complexes from the RCSB Protein Data Bank, built for retrieval-augmented structure-based drug discovery. Every component of the release is commercially redistributable end-to-end. Version history Date Change 2026-04-15 Initial release: FAISS HNSW index, ESM-2 pocket embeddings, Morgan ligand embeddings, HDF5 coordinate files, BindingDB staff-curated affinities.… See the full description on the dataset page: https://huggingface.co/datasets/HemanthKari/rcsb-ligand-retrieval-v0.1.tabular10K<n<100K1 likes91 downloads5mo agoHugging Face30shuklaved /eris-retrieval-benchmark Eris GPU-Accelerated Semantic Retrieval Challenge Welcome to the Eris GPU-Accelerated Semantic Retrieval Challenge platform. This repository contains the complete benchmark dataset, baseline implementations, evaluation grading infrastructure, and reference solution. 1. Dataset Overview The Eris Challenge evaluates high-performance semantic retrieval models over scientific literature abstracts derived from SciFact. Benchmark Specifications… See the full description on the dataset page: https://huggingface.co/datasets/shuklaved/eris-retrieval-benchmark.tabular1K<n<10K0 likes91 downloads18d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.