datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biosses-sts
BIOSSES
An MTEB dataset
Massive Text Embedding Benchmark
Biomedical Semantic Similarity Estimation.
Task category
t2t
Domains
Medical
Reference
https://tabilab.cmpe.boun.edu.tr/BIOSSES/DataSet.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["BIOSSES"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biosses-sts.bio-mqm-datasetThis dataset is compiled from the official Amazon repository (all respective licensing applies).
It contains system translations, multiple references, and their quality evaluation on the MQM scale. It accompanies the ACL 2024 paper Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains.
Watch a brief 4 minutes-long video.
Abstract: We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/bio-mqm-dataset.CVPR-BiomedSegFMThis repository contains the BiomedSegFM dataset, a crucial resource for the CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation.
Foundation Models for Interactive 3D Biomedical Image Segmentation (Homepage)
Foundation Models for Text-guided 3D Biomedical Image Segmentation (Homepage)
CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation
Highly recommend watching the webinar recording to learn about the task settings and… See the full description on the dataset page: https://huggingface.co/datasets/junma/CVPR-BiomedSegFM.biospherebiomedica_webdataset_24M
Dataset Card for Dataset Name
Arxiv: Arxiv
|
Website: Biomedica
|
Training instructions: OpenCLIP
|
Tutorial: Google Colab
BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.BioMysteryBench-preview
BioMysteryBench (preview)
A 5-problem preview of BioMysteryBench,
a bioinformatics research benchmark created by Anthropic. Each problem
provides anonymized biological data files and asks a question that requires
real analysis to answer — the source dataset cannot be looked up.
v11 (2026-07): preview refreshed — hb022 and hb053 were removed from the
benchmark; hb024 and hb035 replace them here. See CHANGELOG.md.
Contents
problems.csv / problems.parquet — one row… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/BioMysteryBench-preview.MedThinkVQA
MedThinkVQA
MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning.
Links
GitHub: https://github.com/benluwang/MedThinkVQA
Leaderboard: https://benluwang.github.io/MedThinkVQA/
Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.bias_in_bios
Bias in Bios
Bias in Bios was created by (De-Artega et al., 2019) and published under the MIT license (https://github.com/microsoft/biosbias). The dataset is used to investigate bias in NLP models. It consists of textual biographies used to predict professional occupations, the sensitive attribute is the gender (binary).
The version shared here is the version proposed by (Ravgofel et al., 2020) which slightly smaller due to the unavailability of 5,557 biographies.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LabHC/bias_in_bios.bio-mcp-data
Bio-MCP-Data
A repository containing biological datasets that will be used by BIO-MCP MCP (Model Context Protocol) standard.
About
This repository hosts biological data assets formatted to be compatible with the Model Context Protocol, enabling AI models to efficiently access and process biological information. The data is managed using Git Large File Storage (LFS) to handle large biological datasets.
Purpose
Provide standardized biological datasets for AI… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/bio-mcp-data.bionotes-storagebiomedical_lectures_v2
Vidore Benchmark 2 - MIT Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_v2.mst_geBioDCASE2026_Bird_Counting
BioDCASE 2026 — Bird Counting (Task 6)
Development and evaluation dataset for the Bird Counting task of the BioDCASE 2026 Challenge.
📢 Evaluation set released on 1 June 2026. 10 new held-out aviaries (~380,000 audio files) are now live under eval_aviary_1/ through eval_aviary_10/. See the Evaluation set section below.
Task overview
Estimating the number of individual birds from acoustic recordings is a fundamental challenge in biodiversity monitoring. This task… See the full description on the dataset page: https://huggingface.co/datasets/Emreargin/BioDCASE2026_Bird_Counting.biorxiv-clustering-p2p
BiorxivClusteringP2P.v2
An MTEB dataset
Massive Text Embedding Benchmark
Clustering of titles+abstract from biorxiv across 26 categories.
Task category
t2c
Domains
Academic, Written
Reference
https://api.biorxiv.org/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["BiorxivClusteringP2P.v2"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biorxiv-clustering-p2p.BioMysteryBench-full
BioMysteryBench (full set)
90 mystery-bioinformatics problems. Each problem provides anonymized
biological data files and asks a question that requires real analysis
(alignment, expression, variant calling, motif discovery, structure, etc.)
to answer — the source dataset cannot be looked up.
v11 (2026-07): 9 problems removed and 24 problems edited after an
answer-key audit — see CHANGELOG.md.
Contents
problems.csv / problems.parquet — one row per problem:
id —… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/BioMysteryBench-full.BioRel
Dataset Card for BioRel
Dataset Summary
BioRel Dataset Summary:
BioRel is a comprehensive dataset designed for biomedical relation extraction, leveraging the vast amount of electronic biomedical literature available.
Developed using the Unified Medical Language System (UMLS) as a knowledge base and Medline articles as a corpus, BioRel utilizes Metamap for entity identification and linking, and employs distant supervision for relation labeling.
The training set… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/BioRel.trustworthy-biology-agents-traces
Trustworthy Biology Agents — Run Traces
Raw execution traces from 1,329 agent runs across three coding agents on three
biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed
trace bundle for the study in
manu-tej/ai-scientists; the write-up
lives in that repo's RESULTS.md.
The motivating question is not only whether an agent reaches the right answer, but
whether it behaves like a trustworthy analyst when the task is ambiguous,
under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.pop-1kgp
1000 Genomes high-coverage GRCh38 population panel
A chromosome-sharded PLINK 2 representation of the 3,202-sample 1000 Genomes
high-coverage GRCh38 callset. It contains 73,759,911 variant records across
chr1-chr22, chrX, chrY, and chrMT.
The 75 published payloads occupy about 5.5 GiB. Exact source files,
derivations, metadata corrections, and fidelity boundaries are recorded in
data/artifact.yaml.
Data layout
Each chromosome is one matching PLINK prefix directly… See the full description on the dataset page: https://huggingface.co/datasets/prescience-bio/pop-1kgp.voxpopuli_da_precomputed_17bagieval-gaokao-biology
Dataset Card for "agieval-gaokao-biology"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub.
This dataset contains the contents of the Gaokao Biology subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 .
Citation:
@misc{zhong2023agieval,
title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-biology.BiomniBench-DA
BiomniBench-DA
BiomniBench-DA is the data-analysis instantiation of BiomniBench, a process-level evaluation framework for LLM agents on real-world biomedical research tasks. Each task is a multi-step data analysis derived from a high-impact biomedical publication; agents are graded on the full analytical trajectory against an expert-authored rubric, not only the final answer.
This repository releases 50 of the 100 BiomniBench-DA tasks; the remaining 50 are held out as a private… See the full description on the dataset page: https://huggingface.co/datasets/phylobio/BiomniBench-DA.wikipedia-biology
Dataset Card for wikipedia-biology
Dataset Summary
The dataset consists of text from 87045 Wikipedia articles created by processing all articles in the Wikipedia categories Branches of biology, Biological concepts, Eukaryote biology and Biology terminology, as well as their subcategories recursively till a depth of 4. It was originally created for the purpose of unlearning the domain of biology, although it may be used for other purposes such as biology fine-tuning.
It… See the full description on the dataset page: https://huggingface.co/datasets/jd5697/wikipedia-biology.Medical_Segmentation_Decathlon
🏆 Medical Segmentation Decathlon Dataset
📝 Overview
The Medical Segmentation Decathlon (MSD) is a comprehensive benchmark dataset for validating algorithms in 3D medical image segmentation. It includes 10 distinct tasks, each with unique challenges like small data sizes, unbalanced labels, varying object scales, multi-class labels, and multimodal imaging.
🔗 Dataset Access
🌐 Website: Medical Decathlon
📂 Google Drive: MSD Google Drive
🧩 Task… See the full description on the dataset page: https://huggingface.co/datasets/Novel-BioMedAI/Medical_Segmentation_Decathlon.vepyr_116_GRCh38_ensembl
vepyr cache — Ensembl VEP 116, GRCh38 (Ensembl)
A Parquet conversion of the Ensembl VEP 116 homo_sapiens_ensembl cache for GRCh38, for
use with vepyr, a Rust/DataFusion variant annotation
engine that reproduces VEP's consequence calls. It replaces VEP's Perl-serialised,
gzipped cache files with columnar Parquet that DuckDB, Polars, DataFusion or Spark can read
directly.
Ensembl/GENCODE transcripts only. This is the default VEP cache flavour. The equivalent VEP invocation uses… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_ensembl.wiki_bioThis dataset gathers 728,321 biographies from wikipedia. It aims at evaluating text generation
algorithms. For each article, we provide the first paragraph and the infobox (both tokenized).
For each article, we extracted the first paragraph (text), the infobox (structured data). Each
infobox is encoded as a list of (field name, field value) pairs. We used Stanford CoreNLP
(http://stanfordnlp.github.io/CoreNLP/) to preprocess the data, i.e. we broke the text into
sentences and tokenized both the text and the field values. The dataset was randomly split in
three subsets train (80%), valid (10%), test (10%).biosProfNER_corpus_NER
Description
Gold standard annotations for profession detection in Spanish COVID-19 tweets
The entire corpus contains 10,000 annotated tweets. It has been split into training, validation, and test (60-20-20). The current version contains the training and development set of the shared task with Gold Standard annotations. In addition, it contains the unannotated test, and background sets will be released.
For Named Entity Recognition, profession detection, annotations are distributed… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/ProfNER_corpus_NER.SPACCC_Tokenizer
The Tokenizer for Clinical Cases Written in Spanish
Introduction
This repository contains the tokenization model trained using the SPACCC_TOKEN corpus (https://github.com/PlanTL-SANIDAD/SPACCC_TOKEN). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to tokenize biomedical documents, specially clinical cases written in Spanish.
This model was created using the Apache… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Tokenizer.Bioinformatics
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Bioinformatics.
