datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.MedThinkVQA
MedThinkVQA
MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning.
Links
GitHub: https://github.com/benluwang/MedThinkVQA
Leaderboard: https://benluwang.github.io/MedThinkVQA/
Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.BiomniBench-DA
BiomniBench-DA
BiomniBench-DA is the data-analysis instantiation of BiomniBench, a process-level evaluation framework for LLM agents on real-world biomedical research tasks. Each task is a multi-step data analysis derived from a high-impact biomedical publication; agents are graded on the full analytical trajectory against an expert-authored rubric, not only the final answer.
This repository releases 50 of the 100 BiomniBench-DA tasks; the remaining 50 are held out as a private… See the full description on the dataset page: https://huggingface.co/datasets/phylobio/BiomniBench-DA.biorXiv-pdf
BiorXiv Pdf
BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets.
BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.biodiversity_heritage_library
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 42 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.AIBio-biostatistician-sft
AIBio Biostatistician SFT dataset
200 instruction/response examples (authored with AI support and validated by the creator) that train an AI biostatistician
(AIBio) to coach a researcher through study design — from a first research idea to
a reporting-guideline-aligned protocol.
This dataset is the training source for the
kcheung-inno/AIBio-LoRA adapter
(qLoRA fine-tune of Qwen/Qwen2.5-3B-Instruct).
Contents
File
Description
examples.json
182 workflow… See the full description on the dataset page: https://huggingface.co/datasets/kcheung-inno/AIBio-biostatistician-sft.biodiversity_heritage_library_filtered
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 15 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.BioMatrix-SFT
BioMatrix-SFT
This is the supervised fine-tuning (SFT) / instruction-tuning corpus used to train BioMatrix, a multimodal foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture.
📄 Paper: BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language
💻 Code: https://github.com/QizhiPei/BioMatrix… See the full description on the dataset page: https://huggingface.co/datasets/QizhiPei/BioMatrix-SFT.TransCorpus-bio
TransCorpus-bio
TransCorpus-bio is a large-scale, parallel biomedical corpus consisting of PubMed abstracts (title + abstract), translated with the TransCorpus Toolkit using NLLB-200. It is designed to enable high-quality multi-lingual biomedical language modeling and downstream NLP research.
This dataset was restructured from five separate single-language repositories into one dataset with a config (tab in the dataset viewer) per language, and with each row carrying its source… See the full description on the dataset page: https://huggingface.co/datasets/jknafou/TransCorpus-bio.BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset.
BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers.
Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT.
Corpus statistics:
Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.biographies
📚 Synthetic Biographies
Synthetic Biographies is a dataset designed to facilitate research in factual recall and representation learning in language models. It comprises synthetic biographies of fictional individuals, each associated with sampled attributes like birthplace, university, and employer. The dataset is intended to support training and evaluating small language models (LLMs), particularly in their ability to store and extract factual knowledge.
🧾 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alex-karev/biographies.bioleaflets-biomedical-ner
Dataset Card for BioLeaflets Dataset
Dataset Summary
BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website.
Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately.
This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.biolangThis dataset is based on abstracts from the open access section of EuropePubMed Central to train language models in the domain of biology.bioinstruct
Dataset Card for BioInstruct
GitHub repo: https://github.com/bio-nlp/BioInstruct
Dataset Summary
BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023.
This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better.
Improvements of Llama on 9 common BioMedical tasks are shown in the result section.
Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.BioMedical_Corpus_1024_2040
PMC 1024-2040 Biomedical Fine-Tuning Corpus
Summary
This is a cleaned biomedical long-text corpus for autoregressive language-model
fine-tuning, held-out evaluation, and membership-inference experiments.
split
rows
role
train
10,000
fine-tuning (membership-positive population)
test
1,000
held-out (membership-negative population)
evaluation
700
balanced membership-inference set: 350 members and 350 non-members
The test split is the full 1… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/BioMedical_Corpus_1024_2040.BiochemForge
BiochemForge
BiochemForge is a provenance-first biology, chemistry, and biochemistry post-training mixture for
mechanistic explanation, quantitative derivation, experimental inference, and consistency between
reasoning and final answers.
Dataset summary
Slice
Records
Purpose
SFT train
99,773
Supervised post-training
SFT validation
2,052
Model selection and early stopping
SFT test
1,093
Internal held-out evaluation
Solver-verified records
27,657… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/BiochemForge.bio-safety-peft-lora
CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset
This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct).
The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios.
🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.BioPhys-Bridge
BioPhys-Bridge is a physics-grounded scientific reasoning dataset for AI-for-Science agents. The dataset subtype, Sci-Evo, represents each record as a Physics-Grounded Scientific Evolution Case linking:
physical model -> quantitative evidence -> biological mechanism -> agent decision
Each case is built from open-access scientific literature and includes evidence-linked text, tables, formulas, figure/caption blocks, normalized quantitative measurements, biophysical model fields, biological… See the full description on the dataset page: https://huggingface.co/datasets/qyxu1994/BioPhys-Bridge.bcs-biostatistics-study
BCS Medical Dataset — biostatistika
Medicinski studijski materijal na bosanskom/hrvatskom/srpskom, obrađen automatizovanim
inbox pipeline-om (ekstrakcija, OCR, chunking, AI generacija s determinističkom validacijom).
Struktura
Fajl
Sadržaj
ispitna.jsonl
postojeća ispitna pitanja (stari testovi/zbornici): question, options, answer
qna.jsonl
AI-generirani QnA parovi (validacija V1-V4)
flashcards.jsonl / flashcards.csv
kartice front/back za učenje… See the full description on the dataset page: https://huggingface.co/datasets/adobug/bcs-biostatistics-study.ALIA-es-biomedical-synthetic-instructions
Dataset Introduction
The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision.
It contains:
639,456 instances
961,073,205 tokens
14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.BioProBench
BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning
BioProBench is the first large-scale, integrated multi-task benchmark for biological protocol understanding and reasoning, specifically designed for large language models (LLMs). It moves beyond simple QA to encompass a comprehensive suite of tasks critical for procedural text comprehension.
Biological protocols are the fundamental bedrock of reproducible and safe life… See the full description on the dataset page: https://huggingface.co/datasets/bowenxian/BioProBench.biographical
Biographical Dataset for Relation Extraction (RE)
Overview
This dataset is a reconstructed version of the Biographical Dataset, specifically designed for relation extraction (RE) tasks. It serves as a valuable resource for digital humanities (DH) and historical research, enabling the study of relationships within biographical data. The dataset is generated by automatically aligning sentences from Wikipedia articles with structured data sourced from platforms like… See the full description on the dataset page: https://huggingface.co/datasets/Despina/biographical.task686_mmmlu_answer_generation_college_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task686_mmmlu_answer_generation_college_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task686_mmmlu_answer_generation_college_biology.GMASS-probe-set-v1.0
MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages
Project Summary
We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.BioTool
BioTool
BioTool is a large-scale, function-calling benchmark and training corpus for the
biomedical domain. It pairs natural-language biomedical questions with the
correct tool call (function name + JSON arguments) that answers them, drawn
from 127 tools spanning the three flagship public APIs:
NCBI E-utilities (einfo, esearch, esummary, efetch, elink, ecitmatch) plus BLAST
UniProt REST (uniprotkb, uniref, uniparc, proteomes, taxonomy, keywords, human_diseases, …)
Ensembl REST… See the full description on the dataset page: https://huggingface.co/datasets/gxx27/BioTool.fineweb-edu-bio
FineWeb-Edu Bio
This is a paragraph-level subset of LeoZotos/fineweb-edu-topics ranked by
biopsychology_similarity. The 2.5B configuration is the highest-ranked
core. The 5B configuration contains that same core plus the extension; the
shared core files are stored only once.
Token budgets use allenai/OLMo-2-0425-1B at revision
stage1-step1907359-tokens4001B and include one EOS document boundary per
paragraph. The paragraph crossing each target is retained, so the actual token… See the full description on the dataset page: https://huggingface.co/datasets/LeoZotos/fineweb-edu-bio.BioManufacturingBench
BioManufacturingBench v1.0.0
BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded
biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation,
process diagnosis, microscopy count-range estimation, strict output formatting, and
abstention. Every primary score is computed by a deterministic rule; no score uses an
LLM judge. Public records are deliberately answer-free so the benchmark remains useful
for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.task699_mmmlu_answer_generation_high_school_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task699_mmmlu_answer_generation_high_school_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task699_mmmlu_answer_generation_high_school_biology.adaption-science-biochem-seed
Science Q&A
Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers.
Rows
12,458
Domain
science
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-biochem-seed.k-12-biology-60k
11-12-biology-60k
Synthetic, original expository text aligned to the CBSE/NCERT-style Class 11-12 Biology curriculum.
Rows: 59,977
Total words: 63,821,097
Subject(s): Science (Biology, Classes 11-12)
Rows per grade: 11: 34,080, 12: 25,897
Fields
Field
Type
Description
text
string
The generated passage
subject
string
Subject name
grade
int
Grade level
word_count
int
Number of words in text
