Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01camel-ai /biology CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs. We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.texttext-generation10K<n<100K58 likes7.2k downloads3y agoHugging Face02bio-nlp-umass /MedThinkVQA MedThinkVQA MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning. Links GitHub: https://github.com/benluwang/MedThinkVQA Leaderboard: https://benluwang.github.io/MedThinkVQA/ Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.imagequestion-answering1K<n<10K12 likes5.1k downloads5mo agoHugging Face03phylobio /BiomniBench-DAgated BiomniBench-DA BiomniBench-DA is the data-analysis instantiation of BiomniBench, a process-level evaluation framework for LLM agents on real-world biomedical research tasks. Each task is a multi-step data analysis derived from a high-impact biomedical publication; agents are graded on the full analytical trajectory against an expert-authored rubric, not only the final answer. This repository releases 50 of the 100 BiomniBench-DA tasks; the remaining 50 are held out as a private… See the full description on the dataset page: https://huggingface.co/datasets/phylobio/BiomniBench-DA.text-generationn<1K26 likes2k downloads4mo agoHugging Face04laion /biorXiv-pdf BiorXiv Pdf BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets. BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.documentfeature-extraction1K<n<10K4 likes1.3k downloads2y agoHugging Face05common-pile /biodiversity_heritage_library Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 42 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.texttext-generation10M<n<100M2 likes1k downloads1y agoHugging Face06kcheung-inno /AIBio-biostatistician-sft AIBio Biostatistician SFT dataset 200 instruction/response examples (authored with AI support and validated by the creator) that train an AI biostatistician (AIBio) to coach a researcher through study design — from a first research idea to a reporting-guideline-aligned protocol. This dataset is the training source for the kcheung-inno/AIBio-LoRA adapter (qLoRA fine-tune of Qwen/Qwen2.5-3B-Instruct). Contents File Description examples.json 182 workflow… See the full description on the dataset page: https://huggingface.co/datasets/kcheung-inno/AIBio-biostatistician-sft.texttext-generationn<1K0 likes593 downloads6d agoHugging Face07common-pile /biodiversity_heritage_library_filtered Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 15 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.texttext-generation10M<n<100M2 likes585 downloads1y agoHugging Face08QizhiPei /BioMatrix-SFT BioMatrix-SFT This is the supervised fine-tuning (SFT) / instruction-tuning corpus used to train BioMatrix, a multimodal foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture. 📄 Paper: BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language 💻 Code: https://github.com/QizhiPei/BioMatrix… See the full description on the dataset page: https://huggingface.co/datasets/QizhiPei/BioMatrix-SFT.texttext-generation10M<n<100M1 likes575 downloads4mo agoHugging Face09jknafou /TransCorpus-bio TransCorpus-bio TransCorpus-bio is a large-scale, parallel biomedical corpus consisting of PubMed abstracts (title + abstract), translated with the TransCorpus Toolkit using NLLB-200. It is designed to enable high-quality multi-lingual biomedical language modeling and downstream NLP research. This dataset was restructured from five separate single-language repositories into one dataset with a config (tab in the dataset viewer) per language, and with each row carrying its source… See the full description on the dataset page: https://huggingface.co/datasets/jknafou/TransCorpus-bio.texttranslation100M<n<1B0 likes532 downloads17d agoHugging Face10IVN-RIN /BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset. BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers. Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT. Corpus statistics: Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.texttext-generation10M<n<100M7 likes257 downloads2y agoHugging Face11alex-karev /biographies 📚 Synthetic Biographies Synthetic Biographies is a dataset designed to facilitate research in factual recall and representation learning in language models. It comprises synthetic biographies of fictional individuals, each associated with sampled attributes like birthplace, university, and employer. The dataset is intended to support training and evaluating small language models (LLMs), particularly in their ability to store and extract factual knowledge. 🧾 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alex-karev/biographies.texttext-generation100K<n<1M4 likes228 downloads1y agoHugging Face12ruslan /bioleaflets-biomedical-ner Dataset Card for BioLeaflets Dataset Dataset Summary BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website. Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately. This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.texttext-generation1K<n<10K4 likes223 downloads4y agoHugging Face13EMBO /biolangThis dataset is based on abstracts from the open access section of EuropePubMed Central to train language models in the domain of biology.text-generation2 likes171 downloads4y agoHugging Face14bio-nlp-umass /bioinstruct Dataset Card for BioInstruct GitHub repo: https://github.com/bio-nlp/BioInstruct Dataset Summary BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023. This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better. Improvements of Llama on 9 common BioMedical tasks are shown in the result section. Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.texttext-generation10K<n<100K25 likes144 downloads2y agoHugging Face15spadeMIA /BioMedical_Corpus_1024_2040 PMC 1024-2040 Biomedical Fine-Tuning Corpus Summary This is a cleaned biomedical long-text corpus for autoregressive language-model fine-tuning, held-out evaluation, and membership-inference experiments. split rows role train 10,000 fine-tuning (membership-positive population) test 1,000 held-out (membership-negative population) evaluation 700 balanced membership-inference set: 350 members and 350 non-members The test split is the full 1… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/BioMedical_Corpus_1024_2040.texttext-generation10K<n<100K0 likes140 downloads3d agoHugging Face160xKitkat /BiochemForge BiochemForge BiochemForge is a provenance-first biology, chemistry, and biochemistry post-training mixture for mechanistic explanation, quantitative derivation, experimental inference, and consistency between reasoning and final answers. Dataset summary Slice Records Purpose SFT train 99,773 Supervised post-training SFT validation 2,052 Model selection and early stopping SFT test 1,093 Internal held-out evaluation Solver-verified records 27,657… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/BiochemForge.tabularquestion-answering100K<n<1M1 likes115 downloads2mo agoHugging Face17devsgnr /bio-safety-peft-lora CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct). The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios. 🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.texttext-generation1K<n<10K0 likes113 downloads26d agoHugging Face18qyxu1994 /BioPhys-Bridge BioPhys-Bridge is a physics-grounded scientific reasoning dataset for AI-for-Science agents. The dataset subtype, Sci-Evo, represents each record as a Physics-Grounded Scientific Evolution Case linking: physical model -> quantitative evidence -> biological mechanism -> agent decision Each case is built from open-access scientific literature and includes evidence-linked text, tables, formulas, figure/caption blocks, normalized quantitative measurements, biophysical model fields, biological… See the full description on the dataset page: https://huggingface.co/datasets/qyxu1994/BioPhys-Bridge.textquestion-answeringn<1K1 likes95 downloads4mo agoHugging Face19adobug /bcs-biostatistics-study BCS Medical Dataset — biostatistika Medicinski studijski materijal na bosanskom/hrvatskom/srpskom, obrađen automatizovanim inbox pipeline-om (ekstrakcija, OCR, chunking, AI generacija s determinističkom validacijom). Struktura Fajl Sadržaj ispitna.jsonl postojeća ispitna pitanja (stari testovi/zbornici): question, options, answer qna.jsonl AI-generirani QnA parovi (validacija V1-V4) flashcards.jsonl / flashcards.csv kartice front/back za učenje… See the full description on the dataset page: https://huggingface.co/datasets/adobug/bcs-biostatistics-study.textquestion-answeringn<1K0 likes93 downloads17d agoHugging Face20SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes91 downloads5mo agoHugging Face21bowenxian /BioProBench BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning BioProBench is the first large-scale, integrated multi-task benchmark for biological protocol understanding and reasoning, specifically designed for large language models (LLMs). It moves beyond simple QA to encompass a comprehensive suite of tasks critical for procedural text comprehension. Biological protocols are the fundamental bedrock of reproducible and safe life… See the full description on the dataset page: https://huggingface.co/datasets/bowenxian/BioProBench.texttext-generation1K<n<10K0 likes81 downloads9mo agoHugging Face22Despina /biographical Biographical Dataset for Relation Extraction (RE) Overview This dataset is a reconstructed version of the Biographical Dataset, specifically designed for relation extraction (RE) tasks. It serves as a valuable resource for digital humanities (DH) and historical research, enabling the study of relationships within biographical data. The dataset is generated by automatically aligning sentences from Wikipedia articles with structured data sourced from platforms like… See the full description on the dataset page: https://huggingface.co/datasets/Despina/biographical.textfeature-extraction1M<n<10M0 likes77 downloads3mo agoHugging Face23Lots-of-LoRAs /task686_mmmlu_answer_generation_college_biology Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task686_mmmlu_answer_generation_college_biology Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task686_mmmlu_answer_generation_college_biology.texttext-generationn<1K0 likes76 downloads2y agoHugging Face24BioinstLab /GMASS-probe-set-v1.0 MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages Project Summary We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.texttext-generationn<1K1 likes71 downloads15d agoHugging Face25gxx27 /BioTool BioTool BioTool is a large-scale, function-calling benchmark and training corpus for the biomedical domain. It pairs natural-language biomedical questions with the correct tool call (function name + JSON arguments) that answers them, drawn from 127 tools spanning the three flagship public APIs: NCBI E-utilities (einfo, esearch, esummary, efetch, elink, ecitmatch) plus BLAST UniProt REST (uniprotkb, uniref, uniparc, proteomes, taxonomy, keywords, human_diseases, …) Ensembl REST… See the full description on the dataset page: https://huggingface.co/datasets/gxx27/BioTool.textquestion-answering1K<n<10K1 likes69 downloads5mo agoHugging Face26LeoZotos /fineweb-edu-bio FineWeb-Edu Bio This is a paragraph-level subset of LeoZotos/fineweb-edu-topics ranked by biopsychology_similarity. The 2.5B configuration is the highest-ranked core. The 5B configuration contains that same core plus the extension; the shared core files are stored only once. Token budgets use allenai/OLMo-2-0425-1B at revision stage1-step1907359-tokens4001B and include one EOS document boundary per paragraph. The paragraph crossing each target is retained, so the actual token… See the full description on the dataset page: https://huggingface.co/datasets/LeoZotos/fineweb-edu-bio.tabulartext-generation10M<n<100M0 likes69 downloads20d agoHugging Face27capicu-ai /BioManufacturingBench BioManufacturingBench v1.0.0 BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation, process diagnosis, microscopy count-range estimation, strict output formatting, and abstention. Every primary score is computed by a deterministic rule; no score uses an LLM judge. Public records are deliberately answer-free so the benchmark remains useful for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.textquestion-answering1K<n<10K0 likes68 downloads2mo agoHugging Face28Lots-of-LoRAs /task699_mmmlu_answer_generation_high_school_biology Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task699_mmmlu_answer_generation_high_school_biology Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task699_mmmlu_answer_generation_high_school_biology.texttext-generationn<1K0 likes67 downloads2y agoHugging Face29rodriguescarson /adaption-science-biochem-seed Science Q&A Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers. Rows 12,458 Domain science Format data.parquet, one row per example Licence apache-2.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-biochem-seed.tabularquestion-answering10K<n<100K0 likes67 downloads15d agoHugging Face30sujitpandey /k-12-biology-60k 11-12-biology-60k Synthetic, original expository text aligned to the CBSE/NCERT-style Class 11-12 Biology curriculum. Rows: 59,977 Total words: 63,821,097 Subject(s): Science (Biology, Classes 11-12) Rows per grade: 11: 34,080, 12: 25,897 Fields Field Type Description text string The generated passage subject string Subject name grade int Grade level word_count int Number of words in text tabulartext-generation10K<n<100K0 likes67 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.