Protein
OpenMed-NER-ProteinDetect-SuperClinical-141MOpenMed-NER-ProteinDetect-SnowMed-568MOpenMed-NER-ProteinDetect-TinyMed-135MOpenMed-NER-ProteinDetect-ElectraMed-560MOpenMed-NER-ProteinDetect-BioMed-109MOpenMed-NER-ProteinDetect-EuroMed-212MOpenMed-NER-ProteinDetect-BioMed-335MOpenMed-NER-ProteinDetect-MultiMed-568M
Datasets
All datasets matching “Protein”claude-protein-binder-design
Claude protein binder design — data release v1.0
1,440 de novo miniprotein binders (50 to 120 residues) against 16 targets, designed by two Claude models operating as autonomous protein-design agents (Mythos Preview, 900 designs; Opus 4.8, 540 designs) and characterized at two contract research organizations, Adaptyv Bio (cell-free expression; SPR/BLI kinetics with the design immobilized) and Twist Bioscience (Fc-fusion expression; capture SPR with a six-point antigen… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/claude-protein-binder-design.ProteinGym_v1
ProteinGym
ProteinGym is a benchmark suite for evaluating protein fitness prediction and design models. It includes both substitution and indel mutations, a wide variety of experimentally assayed proteins, and clinically annotated mutations that are relevant to human disease. In total, ProteinGym includes nearly 3 million different mutations.
Dataset Details
ProteinGym is split into four separate benchmarks, based on the prediction target and the type of mutation… See the full description on the dataset page: https://huggingface.co/datasets/OATML-Markslab/ProteinGym_v1.ProteinGym_v0.1
ProteinGym benchmarks overview
ProteinGym is an extensive set of Deep Mutational Scanning (DMS) assays curated to enable thorough comparisons of various mutation effect predictors indifferent regimes. It is comprised of two benchmarks: 1) a substitution benchmark which consists of the experimental characterisation of ∼1.5M missense variants across 87 DMS assays 2) an indel benchmark that includes ∼300k mutants across 7 DMS assays.
Each processed file in each benchmark corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/OATML-Markslab/ProteinGym_v0.1.group_mpnn
Curated ProteinMPNN training dataset
The multi-chain training data for ProteinMPNN
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets… See the full description on the dataset page: https://huggingface.co/datasets/ProteinMPNN/group_mpnn.protein-docs
Protein Documents (Parquet)
Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata.
Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M
Document Schemes
Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.protein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
