datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.airbnb-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval_beir.pointer-retrievaltweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.opengloss-v2.4-retrieval-pairs
OpenGloss v2.4 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an example paired with its own sense's gloss (positive), and optional sampled cross-headword same-domain negatives. Every pair carries both spans, both reading levels, and live_senses, so a consumer can filter or reweight by… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.4-retrieval-pairs.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13twice_tat_qa_retrieval
TATQA-Retrieval-ko
Constructed a TAT (Textual and Tabular) QA dataset based on information collected from Korean financial reports.
twice_kr_market_report_retrieval
FinMarketReport-Retrieval-ko
Constructed a Retrieval dataset related to the stock market, based on Korean Financial Reports.
twice_kr_news_retrieval
FinNews-Retrieval-ko
Constructed a Retrieval dataset based on Korean financial news articles.
esg_cid_retrieval
Enhancing Retrieval for ESGLLM via ESG-CID -- A Disclosure Content Index Finetuning Dataset for Mapping GRI and ESRS
Usage
from datasets import load_dataset
# document chunks: train/dev/test_gri/test_esrs
documents = load_dataset("esgllm/esg_cid_retrieval", "documents")
# queries (disclosure text): train/dev/test_gri/test_esrs
queries = load_dataset("esgllm/esg_cid_retrieval", "queries")
# training triplets: train/dev
triplets = load_dataset("esgllm/esg_cid_retrieval"… See the full description on the dataset page: https://huggingface.co/datasets/airefinery/esg_cid_retrieval.opengloss-v2.3-retrieval-pairs
Superseded by OpenGloss v2.4 (2026-09-25): every sense now has search queries, QA pairs and verified examples (v2.3 had them only for core and tier 2); level x register definitions and leveled contrasts and explanations are added; and the pretraining corpus no longer contains duplicate documents. v2.3 stays published for reproducibility.
OpenGloss v2.3 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-retrieval-pairs.opengloss-v2.1-retrieval-pairs
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-pairs.treb_table_retrievalcc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15newbalance_shoe_insole_retrieval_and_packing_0611This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 298,
"total_frames": 4967405,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:298"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0611.open-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokensopengloss-v2.2-retrieval-pairs
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-retrieval-pairs.shoe_insole_retrieval_and_packing0515This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 101,
"total_frames": 206299,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/shoe_insole_retrieval_and_packing0515.newbalance_shoe_insole_retrieval_and_packing_0604This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 108,
"total_frames": 1635459,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:108"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0604.mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15tda-gnn-scientific-retrievalPathology-Retrieval-Benchmark
Pathology Retrieval Benchmark — Rerank Cache
Privacy-safe embedding cache for rerank-only reproduction (no document IDs or file paths).
Files
One Parquet file per query: entry_{id}.parquet (100 entries).
Schema
Column
Description
kind
query_text, query_image, or candidate
candidate_idx
0–19 for candidates; null for queries
score
Text-retrieval score (candidates only)
embedding
ColQwen multi-vector embedding (N, 128)
Each file… See the full description on the dataset page: https://huggingface.co/datasets/HZ-VUW/Pathology-Retrieval-Benchmark.opengloss-v2.0-retrieval-pairs
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-retrieval-pairs.mscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01karl-with-retrieval_bm25_v2
Dataset Card for "karl-with-retrieval_bm25_v2"
More Information needed
kbmill-brick-retrieval
KBMill Brick Retrieval Demos
Portable, residual-honest knowledge bricks turned into retrieval evaluation sets for KBMill — the public mill at kbmill.com. Shelf packages live in kbmill-brick-library.
These are not unbounded wiki dumps or synthetic QA. They come from real KBMill manufacturing: bounded packages with muted residual junk filtered where applicable, craft notes, security report in the ZIP, and published, re-runnable cosine retrieval evidence.
Config
Queries… See the full description on the dataset page: https://huggingface.co/datasets/CMiller/kbmill-brick-retrieval.rcsb-ligand-retrieval-v0.1
The RCSB Protein-Ligand Retrieval Index
An openly licensed retrieval dataset of 72,199 protein-ligand
complexes from the RCSB Protein Data Bank, built for retrieval-augmented
structure-based drug discovery. Every component of the release is
commercially redistributable end-to-end.
Version history
Date
Change
2026-04-15
Initial release: FAISS HNSW index, ESM-2 pocket embeddings, Morgan ligand embeddings, HDF5 coordinate files, BindingDB staff-curated affinities.… See the full description on the dataset page: https://huggingface.co/datasets/HemanthKari/rcsb-ligand-retrieval-v0.1.eris-retrieval-benchmark
Eris GPU-Accelerated Semantic Retrieval Challenge
Welcome to the Eris GPU-Accelerated Semantic Retrieval Challenge platform. This repository contains the complete benchmark dataset, baseline implementations, evaluation grading infrastructure, and reference solution.
1. Dataset Overview
The Eris Challenge evaluates high-performance semantic retrieval models over scientific literature abstracts derived from SciFact.
Benchmark Specifications… See the full description on the dataset page: https://huggingface.co/datasets/shuklaved/eris-retrieval-benchmark.
