Team Ai
Modelpublic

OneScience-Group/ESM

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes34downloads
Model Card

<p align="center"> <strong> <span style="font-size: 30px;">ESM</span> </strong> </p>

Model Introduction

ESM (Evolutionary Scale Modeling) is a family of protein language models released by Meta AI / FAIR. It can be used for protein representation extraction, structure prediction, variant effect scoring, and fixed-backbone sequence design.

Paper: Evolutionary-scale prediction of atomic-level protein structure with a language model https://www.science.org/doi/10.1126/science.ade2574

Model Description

This model package provides PyTorch inference support for ESM-1, ESM-2, MSA Transformer, ESMFold, ESM-1v, and ESM-IF1, together with adaptations for running on DCUs. Sample data is distributed with the Hugging Face model repository OneScience-Group/ESM.

Use Cases

ScenarioDescription
Protein representation extractionTakes a FASTA file as input and outputs per-token, mean-pooled, BOS, or contact representations
Protein structure predictionTakes one or more amino acid sequences as input and outputs corresponding PDB structure files
Variant effect scoringTakes a wild-type sequence and a DMS mutation table as input and outputs mutation effect scores
Fixed-backbone sequence designTakes a PDB / CIF structure and chain ID as input and samples candidate sequences that satisfy the backbone constraints
Structure-conditioned sequence scoringTakes a structure and candidate sequences as input and computes their conditional log-likelihoods
Hugging Face / OneCode executionAfter downloading the model project, quickly verifies that the scripts run correctly in a life-sciences runtime environment

Usage Guide

1. OneCode Usage

Try one-click AI4S development in the OneCode online environment:

Try one-click AI4S development

2. Manual Installation and Usage

Hardware Requirements

  • —GPU or DCU is recommended.
  • —A CPU can be used for import checks and lightweight configuration tests; full training and inference will be slow.
  • —DCU users must install DTK in advance. DTK 25.04.2 or later is recommended, or a OneScience-recommended version matching the current cluster.

Environment Check

  • —NVIDIA GPU:
bash
nvidia-smi
  • —Hygon DCU:
bash
hy-smi

Download the Model Package

bash
hf download OneScience-Group/ESM --local-dir ./ESM
cd ESM

This model package includes a small set of sample data that can be used directly to validate the default workflow.

Install the Runtime Environment

DCU Environment

bash
# Activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation supported
pip install onescience[bio-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai
bash
# If required libraries cannot be found, activate the CUDA compatibility environment as follows:
source ${ROCM_PATH}/cuda/env.sh
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH"

After installation, return to the model package directory:

bash
cd ./ESM

Training and Inference Data Overview

The FASTA, PDB / CIF, and DMS files used in the ESM examples are distributed with the Hugging Face model repository OneScience-Group/ESM. After downloading the complete model package, the files are available under data/. This model package does not include a training entry point; the data is intended for example inference and workflow validation. You can also download only the data directory:

bash
hf download OneScience-Group/ESM \
  --include "data/**" \
  --local-dir .

Model Weights

The repository includes multiple ESM model checkpoints under weight/; select the appropriate checkpoint for inference.

Preparing Weights

Place the required ESM weights in the following directory:

text
weight/
  checkpoints/
    esm2_t6_8M_UR50D.pt
    esmfold_3B_v1.pt
    esm1v_t33_650M_UR90S_1.pt
    esm_if1_gvp4_t16_142M_UR50.pt
    ...

When using a shared runtime environment, you can specify the weights location through an environment variable:

bash
export ESM_WEIGHT_DIR=/path/to/esm/weight

The default example uses:

  • —weight/checkpoints/esm2_t6_8M_UR50D.pt

The ESMFold, ESM-1v, and ESM-IF1 examples require their respective weights to be available.

Default Example

bash
bash scripts/infer.sh

The default example reads data/fasta/few_proteins.fasta, extracts protein representations using esm2_t6_8M_UR50D.pt, and saves the results to outputs/embeddings/.

Sequence Representation Extraction

bash
python scripts/extract.py \
  weight/checkpoints/esm2_t6_8M_UR50D.pt \
  data/fasta/few_proteins.fasta \
  outputs/embeddings \
  --include mean per_tok \
  --repr_layers 6

ESMFold Structure Prediction

bash
python scripts/fold.py \
  -i data/fasta/few_proteins.fasta \
  -o outputs/pdb \
  --model-dir weight \
  --cpu-only

The output directory will contain one or more .pdb files. For production GPU / DCU inference, remove --cpu-only and configure --chunk-size or --max-tokens-per-batch according to the available accelerator memory.

You can also explicitly enable ESMFold via the default script:

bash
RUN_ESMFOLD=1 bash scripts/infer.sh

Inverse Folding — Sequence Sampling

bash
python scripts/inverse_folding/sample_sequences.py \
  data/inverse_folding/5YH2.pdb \
  --chain A \
  --outpath outputs/sampled_seqs.fasta \
  --num-samples 1 \
  --nogpu

Inverse Folding — Sequence Scoring

bash
python scripts/inverse_folding/score_log_likelihoods.py \
  data/inverse_folding/5YH2.pdb \
  data/inverse_folding/5YH2_mutated_seqs.fasta \
  --chain A \
  --outpath outputs/sequence_scores.csv \
  --nogpu

Variant Effect Prediction

Variant effect prediction requires a wild-type sequence that is consistent with the mutation annotations in the DMS table:

bash
python scripts/variant_prediction/predict.py \
  --model-location esm1v_t33_650M_UR90S_1 \
  --sequence "${ESM_VARIANT_SEQUENCE}" \
  --dms-input data/variant_prediction/BLAT_ECOLX_Ranganathan2015.csv \
  --mutation-col mutant \
  --dms-output outputs/variant_prediction.csv \
  --offset-idx 24 \
  --scoring-strategy wt-marginals

Data Format

Sample data is stored under data/ by default:

text
data/
  fasta/
    few_proteins.fasta
    some_proteins.fasta
  inverse_folding/
    5YH2.pdb
    5YH2.cif
    5YH2_mutated_seqs.fasta
    example.json
  variant_prediction/
    BLAT_ECOLX_Ranganathan2015.csv
    rho_pp.csv
    aggregated_rho.csv
    aggregated_rho_round3.csv

In this structure:

  • —FASTA files are used for sequence representation extraction and structure prediction.
  • —PDB / CIF files are used for inverse folding sampling and structure-conditioned sequence scoring.
  • —Variant effect prediction CSV files must include a mutation column. The default column name is mutant, and mutations use notation such as A123B.
  • —For custom DMS data, the wild-type amino acid at each mutated position in the sequence provided via --sequence must match the corresponding mutation annotation.

Verification

Static import check:

bash
python scripts/check_import_boundaries.py

Syntax check:

bash
python -B -c "import ast, pathlib; [ast.parse(p.read_text(encoding='utf-8'), filename=str(p)) for root in ['model', 'scripts', 'tests'] for p in pathlib.Path(root).rglob('*.py')]"

OneScience Official Information

PlatformOneScience Main RepositorySkills Repository
Giteehttps://gitee.com/onescience-ai/onesciencehttps://gitee.com/onescience-ai/oneskills
GitHubhttps://github.com/onescience-ai/OneSciencehttps://github.com/onescience-ai/oneskills

Citation & License

  • —This repository is adapted from the open-source ESM model to support DCUs.
  • —The ESM source code is licensed under the MIT License; see LICENSE. For the usage terms governing model weights and data, refer to the documentation provided by the respective publishers.
  • —For scientific use, please cite the corresponding original ESM paper for each submodel used. For ESM-2 / ESMFold, cite: Evolutionary-scale prediction of atomic-level protein structure with a language model.