bioscience
mf-pcba-bind
MF-PCBA-Bind
Code for generating these datasets can be found at https://github.com/Leash-Labs/mf-pcba-bind
Protein-ligand binding prediction datasets derived from the MF-PCBA benchmark.
This repository extends the original MF-PCBA dataset by:
Filtering to binding assays only (excluding phenotypic assays)
Removing PAINS (pan-assay interference compounds) that show non-specific activity
Providing pre-built validation and test splits for protein-ligand binding prediction
Aggregating… See the full description on the dataset page: https://huggingface.co/datasets/Leash-Biosciences/mf-pcba-bind.biosciences-competency-questions-sample
Open Biosciences Competency Questions (Sample)
Dataset Description
A curated collection of 15 competency questions (CQs) for evaluating and guiding knowledge graph construction in biosciences research. Each question includes structured entities with standardized CURIEs, gold-standard knowledge graph paths using BioLink predicates, executable multi-API workflow steps, and source provenance.
What Are Competency Questions?
Competency questions are natural language… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-competency-questions-sample.biosciences-sources
Biosciences RAG Source Documents
Dataset Description
This dataset contains 140 page-level document chunks extracted from 10 biomedical research papers. The documents form the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on biosciences topics including knowledge graphs, LLM applications in biomedicine, and protein interaction databases.
Dataset Summary
Total Documents: 140 pages from 10 research papers
Domain: Biomedical NLP… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-sources.papyrus-decoy-eval
Papyrus Decoy Evaluation Set
A protein-ligand binding benchmark dataset derived from the Papyrus bioactivity database, augmented with property-matched decoy molecules for evaluating virtual screening and binding prediction models.
What is this dataset?
This dataset pairs known protein-ligand binders (actives) with synthetically selected decoy molecules — compounds that are physicochemically similar to the actives but are assumed to be non-binders. This setup enables… See the full description on the dataset page: https://huggingface.co/datasets/Leash-Biosciences/papyrus-decoy-eval.biosciences-evaluation-inputs
Biosciences RAG Evaluation Inputs
Dataset Description
This dataset contains RAG inference outputs from 4 retrieval strategies evaluated on 12 biosciences research questions. Each retriever was tested on the same golden testset, producing 48 total records with retrieved contexts and LLM-generated answers ready for RAGAS evaluation.
Dataset Summary
Total Examples: 48 records (12 questions x 4 retrievers)
Retrievers Compared:
Naive — Dense vector similarity… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-evaluation-inputs.biosciences-golden-testset
Biosciences RAG Golden Test Set
Dataset Description
This dataset contains 12 question-answering pairs for evaluating RAG systems on biomedical research topics. The QA pairs were synthetically generated using the RAGAS framework from 140 source documents spanning knowledge graphs, LLM applications in biomedicine, protein interaction databases, and gene-to-phenotype mapping.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation ground truth… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-golden-testset.
