Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /biosses-sts BIOSSES An MTEB dataset Massive Text Embedding Benchmark Biomedical Semantic Similarity Estimation. Task category t2t Domains Medical Reference https://tabilab.cmpe.boun.edu.tr/BIOSSES/DataSet.html How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["BIOSSES"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model) To learn more… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biosses-sts.textsentence-similarityn<1K2 likes22k downloads1y agoHugging Face02zouhar /bio-mqm-datasetThis dataset is compiled from the official Amazon repository (all respective licensing applies). It contains system translations, multiple references, and their quality evaluation on the MQM scale. It accompanies the ACL 2024 paper Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains. Watch a brief 4 minutes-long video. Abstract: We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/bio-mqm-dataset.texttranslation10K<n<100K8 likes16k downloads2y agoHugging Face03junma /CVPR-BiomedSegFMThis repository contains the BiomedSegFM dataset, a crucial resource for the CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation. Foundation Models for Interactive 3D Biomedical Image Segmentation (Homepage) Foundation Models for Text-guided 3D Biomedical Image Segmentation (Homepage) CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation Highly recommend watching the webinar recording to learn about the task settings and… See the full description on the dataset page: https://huggingface.co/datasets/junma/CVPR-BiomedSegFM.3dimage-segmentation24 likes13k downloads7mo agoHugging Face04phuongly84829 /biosphere0 likes9k downloads2h agoHugging Face05BIOMEDICA /biomedica_webdataset_24Mgated Dataset Card for Dataset Name Arxiv: Arxiv &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Website: Biomedica &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Training instructions: OpenCLIP &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Tutorial: Google Colab BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.n>1T40 likes8.4k downloads1mo agoHugging Face06camel-ai /biology CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs. We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.texttext-generation10K<n<100K58 likes7.2k downloads3y agoHugging Face07Anthropic /BioMysteryBench-preview BioMysteryBench (preview) A 5-problem preview of BioMysteryBench, a bioinformatics research benchmark created by Anthropic. Each problem provides anonymized biological data files and asks a question that requires real analysis to answer — the source dataset cannot be looked up. v11 (2026-07): preview refreshed — hb022 and hb053 were removed from the benchmark; hb024 and hb035 replace them here. See CHANGELOG.md. Contents problems.csv / problems.parquet — one row… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/BioMysteryBench-preview.21 likes5.9k downloads3mo agoHugging Face08bio-nlp-umass /MedThinkVQA MedThinkVQA MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning. Links GitHub: https://github.com/benluwang/MedThinkVQA Leaderboard: https://benluwang.github.io/MedThinkVQA/ Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.imagequestion-answering1K<n<10K12 likes5.1k downloads5mo agoHugging Face09LabHC /bias_in_bios Bias in Bios Bias in Bios was created by (De-Artega et al., 2019) and published under the MIT license (https://github.com/microsoft/biosbias). The dataset is used to investigate bias in NLP models. It consists of textual biographies used to predict professional occupations, the sensitive attribute is the gender (binary). The version shared here is the version proposed by (Ravgofel et al., 2020) which slightly smaller due to the unavailability of 5,557 biographies. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LabHC/bias_in_bios.tabulartext-classification100K<n<1M23 likes4.8k downloads3y agoHugging Face10longevity-genie /bio-mcp-data Bio-MCP-Data A repository containing biological datasets that will be used by BIO-MCP MCP (Model Context Protocol) standard. About This repository hosts biological data assets formatted to be compatible with the Model Context Protocol, enabling AI models to efficiently access and process biological information. The data is managed using Git Large File Storage (LFS) to handle large biological datasets. Purpose Provide standardized biological datasets for AI… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/bio-mcp-data.text0 likes4.4k downloads1y agoHugging Face11titasmallick /bionotes-storage0 likes4.1k downloads3h agoHugging Face12vidore /biomedical_lectures_v2 Vidore Benchmark 2 - MIT Dataset (Multilingual) This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions). Dataset Summary The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_v2.imagedocument-question-answering1K<n<10K0 likes3.7k downloads1y agoHugging Face13mania-bioinformatics /mst_gegated0 likes3.6k downloads55m agoHugging Face14Emreargin /BioDCASE2026_Bird_Counting BioDCASE 2026 — Bird Counting (Task 6) Development and evaluation dataset for the Bird Counting task of the BioDCASE 2026 Challenge. 📢 Evaluation set released on 1 June 2026. 10 new held-out aviaries (~380,000 audio files) are now live under eval_aviary_1/ through eval_aviary_10/. See the Evaluation set section below. Task overview Estimating the number of individual birds from acoustic recordings is a fundamental challenge in biodiversity monitoring. This task… See the full description on the dataset page: https://huggingface.co/datasets/Emreargin/BioDCASE2026_Bird_Counting.audioaudio-classification100K<n<1M0 likes3.4k downloads3mo agoHugging Face15mteb /biorxiv-clustering-p2p BiorxivClusteringP2P.v2 An MTEB dataset Massive Text Embedding Benchmark Clustering of titles+abstract from biorxiv across 26 categories. Task category t2c Domains Academic, Written Reference https://api.biorxiv.org/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["BiorxivClusteringP2P.v2"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biorxiv-clustering-p2p.texttext-classification10K<n<100K0 likes3.3k downloads8mo agoHugging Face16Anthropic /BioMysteryBench-fullgated BioMysteryBench (full set) 90 mystery-bioinformatics problems. Each problem provides anonymized biological data files and asks a question that requires real analysis (alignment, expression, variant calling, motif discovery, structure, etc.) to answer — the source dataset cannot be looked up. v11 (2026-07): 9 problems removed and 24 problems edited after an answer-key audit — see CHANGELOG.md. Contents problems.csv / problems.parquet — one row per problem: id —… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/BioMysteryBench-full.66 likes2.8k downloads3mo agoHugging Face17DFKI-SLT /BioRel Dataset Card for BioRel Dataset Summary BioRel Dataset Summary: BioRel is a comprehensive dataset designed for biomedical relation extraction, leveraging the vast amount of electronic biomedical literature available. Developed using the Unified Medical Language System (UMLS) as a knowledge base and Medline articles as a corpus, BioRel utilizes Metamap for entity identification and linking, and employs distant supervision for relation labeling. The training set… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/BioRel.texttext-classification100K<n<1M4 likes2.3k downloads2y agoHugging Face18amanutej /trustworthy-biology-agents-traces Trustworthy Biology Agents — Run Traces Raw execution traces from 1,329 agent runs across three coding agents on three biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed trace bundle for the study in manu-tej/ai-scientists; the write-up lives in that repo's RESULTS.md. The motivating question is not only whether an agent reaches the right answer, but whether it behaves like a trustworthy analyst when the task is ambiguous, under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.tabular1K<n<10K0 likes2.2k downloads3mo agoHugging Face19prescience-bio /pop-1kgp 1000 Genomes high-coverage GRCh38 population panel A chromosome-sharded PLINK 2 representation of the 3,202-sample 1000 Genomes high-coverage GRCh38 callset. It contains 73,759,911 variant records across chr1-chr22, chrX, chrY, and chrMT. The 75 published payloads occupy about 5.5 GiB. Exact source files, derivations, metadata corrections, and fidelity boundaries are recorded in data/artifact.yaml. Data layout Each chromosome is one matching PLINK prefix directly… See the full description on the dataset page: https://huggingface.co/datasets/prescience-bio/pop-1kgp.0 likes2.1k downloads2d agoHugging Face20Biorrith /voxpopuli_da_precomputed_17btabular1M<n<10M0 likes2.1k downloads15d agoHugging Face21hails /agieval-gaokao-biology Dataset Card for "agieval-gaokao-biology" Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub. This dataset contains the contents of the Gaokao Biology subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 . Citation: @misc{zhong2023agieval, title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-biology.textn<1K1 likes2.1k downloads3y agoHugging Face22phylobio /BiomniBench-DAgated BiomniBench-DA BiomniBench-DA is the data-analysis instantiation of BiomniBench, a process-level evaluation framework for LLM agents on real-world biomedical research tasks. Each task is a multi-step data analysis derived from a high-impact biomedical publication; agents are graded on the full analytical trajectory against an expert-authored rubric, not only the final answer. This repository releases 50 of the 100 BiomniBench-DA tasks; the remaining 50 are held out as a private… See the full description on the dataset page: https://huggingface.co/datasets/phylobio/BiomniBench-DA.text-generationn<1K26 likes2k downloads4mo agoHugging Face23jd5697 /wikipedia-biology Dataset Card for wikipedia-biology Dataset Summary The dataset consists of text from 87045 Wikipedia articles created by processing all articles in the Wikipedia categories Branches of biology, Biological concepts, Eukaryote biology and Biology terminology, as well as their subcategories recursively till a depth of 4. It was originally created for the purpose of unlearning the domain of biology, although it may be used for other purposes such as biology fine-tuning. It… See the full description on the dataset page: https://huggingface.co/datasets/jd5697/wikipedia-biology.text10K<n<100K0 likes1.9k downloads2y agoHugging Face24Novel-BioMedAI /Medical_Segmentation_Decathlon 🏆 Medical Segmentation Decathlon Dataset 📝 Overview The Medical Segmentation Decathlon (MSD) is a comprehensive benchmark dataset for validating algorithms in 3D medical image segmentation. It includes 10 distinct tasks, each with unique challenges like small data sizes, unbalanced labels, varying object scales, multi-class labels, and multimodal imaging. 🔗 Dataset Access 🌐 Website: Medical Decathlon 📂 Google Drive: MSD Google Drive 🧩 Task… See the full description on the dataset page: https://huggingface.co/datasets/Novel-BioMedAI/Medical_Segmentation_Decathlon.2 likes1.8k downloads1y agoHugging Face25biodatageeks /vepyr_116_GRCh38_ensembl vepyr cache — Ensembl VEP 116, GRCh38 (Ensembl) A Parquet conversion of the Ensembl VEP 116 homo_sapiens_ensembl cache for GRCh38, for use with vepyr, a Rust/DataFusion variant annotation engine that reproduces VEP's consequence calls. It replaces VEP's Perl-serialised, gzipped cache files with columnar Parquet that DuckDB, Polars, DataFusion or Spark can read directly. Ensembl/GENCODE transcripts only. This is the default VEP cache flavour. The equivalent VEP invocation uses… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_ensembl.tabular1B<n<10B0 likes1.7k downloads4d agoHugging Face26michaelauli /wiki_bioThis dataset gathers 728,321 biographies from wikipedia. It aims at evaluating text generation algorithms. For each article, we provide the first paragraph and the infobox (both tokenized). For each article, we extracted the first paragraph (text), the infobox (structured data). Each infobox is encoded as a list of (field name, field value) pairs. We used Stanford CoreNLP (http://stanfordnlp.github.io/CoreNLP/) to preprocess the data, i.e. we broke the text into sentences and tokenized both the text and the field values. The dataset was randomly split in three subsets train (80%), valid (10%), test (10%).table-to-text100K<n<1M27 likes1.7k downloads3y agoHugging Face27phatpham84453 /bios0 likes1.6k downloads2h agoHugging Face28Biomedical-TeMU /ProfNER_corpus_NER Description Gold standard annotations for profession detection in Spanish COVID-19 tweets The entire corpus contains 10,000 annotated tweets. It has been split into training, validation, and test (60-20-20). The current version contains the training and development set of the shared task with Gold Standard annotations. In addition, it contains the unannotated test, and background sets will be released. For Named Entity Recognition, profession detection, annotations are distributed… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/ProfNER_corpus_NER.text1 likes1.5k downloads5y agoHugging Face29Biomedical-TeMU /SPACCC_Tokenizer The Tokenizer for Clinical Cases Written in Spanish Introduction This repository contains the tokenization model trained using the SPACCC_TOKEN corpus (https://github.com/PlanTL-SANIDAD/SPACCC_TOKEN). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to tokenize biomedical documents, specially clinical cases written in Spanish. This model was created using the Apache… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Tokenizer.text10K<n<100K0 likes1.4k downloads5y agoHugging Face30R2MED /Bioinformatics 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Bioinformatics.texttext-retrieval10K<n<100K2 likes1.4k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.