proteins
Datasets
All datasets matching “proteins”initial-dynamic-proteins
Initial 2,000-protein dataset
This is the canonical local root for the first complete router dataset: 1,000
single-dominant structured-state proteins (label 0) and 1,000 dynamic or
heterogeneous-state proteins (label 1). The fixed split is 1,400 train, 300
validation, and 300 test proteins.
Place Colab's completed ESMFold result files (<sequence_sha256>.npz) in
esmfold_results/. Then import them with:
uv run python scripts/esmfold_dataset.py import
The importer validates every… See the full description on the dataset page: https://huggingface.co/datasets/archiitecture/initial-dynamic-proteins.Proteinscasp14-casp15-cameo-test-proteinsProteinSelfies10 million random examples from Uniref50 representative sequences (October 2023) and computed selfies strings. The strings are stored as input ids from a custom selfies tokenizer. A BERT tokenizer with this vocabulary has been uploaded to this dataset under the files.
You can access the tokenizer like this:
import os
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
repo_path = 'Synthyra/ProteinSelfies'
local_path = 'ProteinSelfies'
files =… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ProteinSelfies.dynamic_proteins_promise_sf_cluster_continuation
Dynamic Protein Benchmarking MVP
This repository contains a first-pass Python package for benchmarking whether protein structure-generation models recover multiple experimentally observed conformations of the same protein.
The MVP targets unconditioned multistate recovery for BioEmu and Boltz across full-MSA, shallow-MSA, and no-MSA style conditions. The initial implementation emphasizes reproducible data structures, evaluation metrics, deterministic outputs, and a notebook UI… See the full description on the dataset page: https://huggingface.co/datasets/archiitecture/dynamic_proteins_promise_sf_cluster_continuation.PROTEINS
Dataset Card for PROTEINS
Dataset Summary
The PROTEINS dataset is a medium molecular property prediction dataset.
Supported Tasks and Leaderboards
PROTEINS should be used for molecular property prediction (aiming to predict whether molecules are enzymes or not), a binary classification task. The score used is accuracy, using a 10-fold cross-validation.
External Use
PyGeometric
To load in PyGeometric, do the following:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/PROTEINS.
