Team Ai
Modelpublic

OneScience-Group/ESM

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes24downloads
README.md286 linesDownload Raw Back to root
1---2license: mit3tasks:4  - protein-structure-prediction5frameworks:6  - pytorch7language:8  - en9  - zh10tags:11  - OneScience12  - Life Sciences13  - Protein Language Model14  - Protein Structure Prediction15  - Variant Effect Prediction16  - ESM17huggingface:18  model: OneScience-Group/ESM19  data_path: data/20---21 22<p align="center">23  <strong>24    <span style="font-size: 30px;">ESM</span>25  </strong>26</p>27 28# Model Introduction29 30ESM (Evolutionary Scale Modeling) is a family of protein language models released by Meta AI / FAIR. It can be used for protein representation extraction, structure prediction, variant effect scoring, and fixed-backbone sequence design.31 32Paper: Evolutionary-scale prediction of atomic-level protein structure with a language model  33https://www.science.org/doi/10.1126/science.ade257434 35# Model Description36 37This model package provides PyTorch inference support for ESM-1, ESM-2, MSA Transformer, ESMFold, ESM-1v, and ESM-IF1, together with adaptations for running on DCUs. Sample data is distributed with the Hugging Face model repository `OneScience-Group/ESM`.38 39# Use Cases40 41| Scenario | Description |42| :---: | :--- |43| Protein representation extraction | Takes a FASTA file as input and outputs per-token, mean-pooled, BOS, or contact representations |44| Protein structure prediction | Takes one or more amino acid sequences as input and outputs corresponding PDB structure files |45| Variant effect scoring | Takes a wild-type sequence and a DMS mutation table as input and outputs mutation effect scores |46| Fixed-backbone sequence design | Takes a PDB / CIF structure and chain ID as input and samples candidate sequences that satisfy the backbone constraints |47| Structure-conditioned sequence scoring | Takes a structure and candidate sequences as input and computes their conditional log-likelihoods |48| Hugging Face / OneCode execution | After downloading the model project, quickly verifies that the scripts run correctly in a life-sciences runtime environment |49 50 51# Usage Guide52 53## 1. OneCode Usage54 55Try one-click AI4S development in the OneCode online environment:56 57[Try one-click AI4S development](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)58 59## 2. Manual Installation and Usage60 61**Hardware Requirements**62 63- GPU or DCU is recommended.64- A CPU can be used for import checks and lightweight configuration tests; full training and inference will be slow.65- DCU users must install DTK in advance. DTK 25.04.2 or later is recommended, or a OneScience-recommended version matching the current cluster.66 67 68 69 70 71**Environment Check**72 73- NVIDIA GPU:74 75```bash76nvidia-smi77```78 79- Hygon DCU:80 81```bash82hy-smi83```84 85### Download the Model Package86 87```bash88hf download OneScience-Group/ESM --local-dir ./ESM89cd ESM90```91 92This model package includes a small set of sample data that can be used directly to validate the default workflow.93 94### Install the Runtime Environment95 96**DCU Environment**97 98```bash99# Activate DTK and CONDA first100conda create -n onescience311 python=3.11 -y101conda activate onescience311102# uv installation supported103pip install onescience[bio-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai104```105 106```bash107# If required libraries cannot be found, activate the CUDA compatibility environment as follows:108source ${ROCM_PATH}/cuda/env.sh109export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"110export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH"111```112 113After installation, return to the model package directory:114 115```bash116cd ./ESM117```118 119### Training and Inference Data Overview120 121The FASTA, PDB / CIF, and DMS files used in the ESM examples are distributed with the [Hugging Face model repository OneScience-Group/ESM](https://huggingface.co/OneScience-Group/ESM). After downloading the complete model package, the files are available under `data/`. This model package does not include a training entry point; the data is intended for example inference and workflow validation. You can also download only the data directory:122 123```bash124hf download OneScience-Group/ESM \125  --include "data/**" \126  --local-dir .127```128### Model Weights129 130The repository includes multiple ESM model checkpoints under `weight/`; select the appropriate checkpoint for inference.131 132### Preparing Weights133 134Place the required ESM weights in the following directory:135 136```text137weight/138  checkpoints/139    esm2_t6_8M_UR50D.pt140    esmfold_3B_v1.pt141    esm1v_t33_650M_UR90S_1.pt142    esm_if1_gvp4_t16_142M_UR50.pt143    ...144```145 146When using a shared runtime environment, you can specify the weights location through an environment variable:147 148```bash149export ESM_WEIGHT_DIR=/path/to/esm/weight150```151 152The default example uses:153 154- `weight/checkpoints/esm2_t6_8M_UR50D.pt`155 156The ESMFold, ESM-1v, and ESM-IF1 examples require their respective weights to be available.157 158### Default Example159 160```bash161bash scripts/infer.sh162```163 164The default example reads `data/fasta/few_proteins.fasta`, extracts protein representations using `esm2_t6_8M_UR50D.pt`, and saves the results to `outputs/embeddings/`.165 166### Sequence Representation Extraction167 168```bash169python scripts/extract.py \170  weight/checkpoints/esm2_t6_8M_UR50D.pt \171  data/fasta/few_proteins.fasta \172  outputs/embeddings \173  --include mean per_tok \174  --repr_layers 6175```176 177### ESMFold Structure Prediction178 179```bash180python scripts/fold.py \181  -i data/fasta/few_proteins.fasta \182  -o outputs/pdb \183  --model-dir weight \184  --cpu-only185```186 187The output directory will contain one or more `.pdb` files. For production GPU / DCU inference, remove `--cpu-only` and configure `--chunk-size` or `--max-tokens-per-batch` according to the available accelerator memory.188 189You can also explicitly enable ESMFold via the default script:190 191```bash192RUN_ESMFOLD=1 bash scripts/infer.sh193```194 195### Inverse Folding — Sequence Sampling196 197```bash198python scripts/inverse_folding/sample_sequences.py \199  data/inverse_folding/5YH2.pdb \200  --chain A \201  --outpath outputs/sampled_seqs.fasta \202  --num-samples 1 \203  --nogpu204```205 206### Inverse Folding — Sequence Scoring207 208```bash209python scripts/inverse_folding/score_log_likelihoods.py \210  data/inverse_folding/5YH2.pdb \211  data/inverse_folding/5YH2_mutated_seqs.fasta \212  --chain A \213  --outpath outputs/sequence_scores.csv \214  --nogpu215```216 217### Variant Effect Prediction218 219Variant effect prediction requires a wild-type sequence that is consistent with the mutation annotations in the DMS table:220 221```bash222python scripts/variant_prediction/predict.py \223  --model-location esm1v_t33_650M_UR90S_1 \224  --sequence "${ESM_VARIANT_SEQUENCE}" \225  --dms-input data/variant_prediction/BLAT_ECOLX_Ranganathan2015.csv \226  --mutation-col mutant \227  --dms-output outputs/variant_prediction.csv \228  --offset-idx 24 \229  --scoring-strategy wt-marginals230```231 232# Data Format233 234Sample data is stored under `data/` by default:235 236```text237data/238  fasta/239    few_proteins.fasta240    some_proteins.fasta241  inverse_folding/242    5YH2.pdb243    5YH2.cif244    5YH2_mutated_seqs.fasta245    example.json246  variant_prediction/247    BLAT_ECOLX_Ranganathan2015.csv248    rho_pp.csv249    aggregated_rho.csv250    aggregated_rho_round3.csv251```252 253In this structure:254 255- FASTA files are used for sequence representation extraction and structure prediction.256- PDB / CIF files are used for inverse folding sampling and structure-conditioned sequence scoring.257- Variant effect prediction CSV files must include a mutation column. The default column name is `mutant`, and mutations use notation such as `A123B`.258- For custom DMS data, the wild-type amino acid at each mutated position in the sequence provided via `--sequence` must match the corresponding mutation annotation.259 260# Verification261 262Static import check:263 264```bash265python scripts/check_import_boundaries.py266```267 268Syntax check:269 270```bash271python -B -c "import ast, pathlib; [ast.parse(p.read_text(encoding='utf-8'), filename=str(p)) for root in ['model', 'scripts', 'tests'] for p in pathlib.Path(root).rglob('*.py')]"272```273 274# OneScience Official Information275 276| Platform | OneScience Main Repository | Skills Repository |277| --- | --- | --- |278| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |279| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |280 281# Citation & License282 283- This repository is adapted from the open-source ESM model to support DCUs.284- The ESM source code is licensed under the MIT License; see `LICENSE`. For the usage terms governing model weights and data, refer to the documentation provided by the respective publishers.285- For scientific use, please cite the corresponding original ESM paper for each submodel used. For ESM-2 / ESMFold, cite: [Evolutionary-scale prediction of atomic-level protein structure with a language model](https://www.science.org/doi/10.1126/science.ade2574).286