OneScience-Group/DiffDock
052
1---2frameworks:3- ""4language:5- en6license: mit7tags:8- OneScience9- protein-ligand molecular docking10- generative11---12<p align="center">13 <strong>14 <span style="font-size: 30px;">DiffDock</span>15 </strong>16</p>17 18# Model Overview19 20DiffDock is a diffusion model for protein-ligand molecular docking proposed by Corso et al. It supports single-complex molecular docking, batch molecular docking, evaluation, and related workflows.21 22# Model Description23 24DiffDock formulates docking as a generative modeling problem. Given a protein receptor and a small-molecule ligand, it jointly models the ligand's translational, rotational, and torsional degrees of freedom through diffusion processes on SO(3)/SE(3), and generates candidate binding conformations.25 26Weights and datasets are not available at the moment. They will be uploaded to Hugging Face soon, and command-line downloads will be supported later.27 28# Use Cases29 30| Use case | Description |31| :---: | :---: |32| Score model training | Takes processed PDBBind or MOAD data as input and outputs a DiffDock score model checkpoint |33| Single-complex molecular docking | Takes a protein PDB file and a ligand SMILES/SDF/MOL2 input, and outputs candidate ligand binding conformations in SDF format |34| Batch molecular docking | Takes a CSV file containing proteins, ligands, and complex names, and outputs sampled conformations in batches |35| Confidence rerank | Uses an additional confidence model to rank sampled conformations |36| Dataset evaluation | Computes metrics such as RMSD for sampling results on a validation or test set, with optional GNINA integration |37 38# Usage39 40## 1. Using OneCode41 42You can try intelligent one-click AI4S programming through the OneCode online environment:43 44[Try intelligent one-click AI4S programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)45 46## 2. Manual Installation and Usage47 48**Hardware Requirements**49 50- Running on a GPU or DCU is recommended.51- CPU can be used for connectivity checks, but it is relatively slow.52- DCU users need to install DTK in advance. DTK 25.04.2 or later is recommended, or the OneScience-recommended version that matches the current cluster.53 54**Software Requirements**55 56Common DiffDock dependencies include PyTorch, PyTorch Geometric, RDKit, OpenBabel, e3nn, torch-scatter, torch-cluster, NumPy, SciPy, tqdm, PyYAML, and others. GNINA energy minimization during dataset evaluation requires installing `gnina` separately.57 58**Environment Checks**59 60- NVIDIA GPU:61 62```bash63nvidia-smi64```65 66- Hygon DCU:67 68```bash69hy-smi70```71 72## Quick Start73 74### 1. Install the Runtime Environment75 76```bash77conda create -n onescience311 python=3.11 -y78conda activate onescience31179pip install onescience[bio] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai80```81 82If the following code cannot find required libraries at runtime, activate CUDA as shown below.83 84```bash85source ${ROCM_PATH}/cuda/env.sh86export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"87export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH"88```89 90### 2. Download the Model Package and Database91 92```bash93# If the model ID on the Hugging Face page uses different capitalization, use the actual published name.94hf download OneScience-Group/diffdock --local-dir ./diffdock95cd diffdock96```97 98### Training Weights and Datasets99 100Weights and datasets are not available at the moment. They will be uploaded to Hugging Face soon, and command-line downloads will be supported later.101 102### 3. Check Files in the Package103 104The current package does not include pretrained weights. The command below should only show source code, configuration files, and example input files:105 106```bash107find . -maxdepth 3 -type f108```109 110If sampling or evaluation is required, first obtain a checkpoint through training, or place external DiffDock score/confidence weights in a local directory. Make sure the weight directory contains `model_parameters.yml`.111 112### 4. Score Model Training113 114Before training, update the data paths in `configs/training.yml` to local paths. Common fields include:115 116- `data.pdbbind_dir`: processed PDBBind data directory.117- `data.moad_dir`: processed MOAD data directory.118- `data.split_train`: list of training-set complexes.119- `data.split_val`: list of validation-set complexes.120- `runtime.log_dir`: training output directory.121 122Start training:123 124```bash125cd scripts126bash train.sh127```128 129After training succeeds, the output directory usually contains:130 131- `model_parameters.yml`: model architecture and training parameters.132- `best_model.pt`: weights with the best validation loss.133- `best_ema_model.pt`: EMA weights.134- `best_inference_epoch_model.pt` or `best_ema_inference_epoch_model.pt`: weights saved when inference validation is enabled.135- `last_model.pt`: training state from the final epoch.136 137### 5. Molecular Docking Sampling138 139Sampling requires a trained score model directory, for example:140 141```text142outputs/train/diffdock_cg_example/143├── model_parameters.yml144└── best_model.pt145```146 147Update the key fields in `configs/sampling.yml` to the actual paths:148 149- `model.model_dir`: score model directory.150- `model.ckpt`: score checkpoint file name.151- `confidence.confidence_model_dir`: confidence model directory; set it to `null` when rerank is not used.152- `input.protein_path`: protein PDB path. You can use `data/6o5u_protein_processed.pdb`.153- `input.ligand_description`: SMILES string or ligand SDF/MOL2 path. You can use `data/6o5u_ligand.sdf`.154- `runtime.out_dir`: sampling output directory.155 156Start single-complex sampling:157 158```bash159cd scripts160bash infer.sh161```162 163### 6. CSV Batch Sampling164 165Set `input.protein_ligand_csv` in `configs/sampling.yml` to the CSV path, and set the single-complex fields to `null` as needed. The CSV is recommended to contain the following columns:166 167```text168complex_name,protein_path,ligand_description,protein_sequence169```170 171Here, `ligand_description` can be either a SMILES string or an SDF/MOL2 file path.172 173### 7. Dataset Evaluation174 175Before evaluation, update the dataset, model, and output directories in `configs/evaluate.yml` to actual paths:176 177```bash178python -m scripts.evaluate --config configs/evaluate.yml179```180 181If `gnina.gnina_minimize=true` is enabled, first make sure the following command can be run directly in the current environment:182 183```bash184gnina --help185```186 187### 8. Confidence Model Training188 189The confidence training entry point is provided by the installed OneScience package. This Hugging Face package does not duplicate that module. After completing score model training, run the following in the OneScience environment:190 191```bash192python -m onescience.confidence.diffdock.confidence_train \193 --original_model_dir outputs/train/diffdock_cg_example \194 --data_dir /path/to/PDBBind_processed \195 --split_train /path/to/splits/timesplit_no_lig_overlap_train \196 --split_val /path/to/splits/timesplit_no_lig_overlap_val197```198 199# Data Preparation200 201DiffDock training usually requires preprocessed PDBBind or MOAD data. It is recommended to organize the data under a unified root directory, for example:202 203```text204${ONESCIENCE_DATASETS_DIR}/diffdock/205├── PDBBind_processed/206├── MOAD_processed/207└── splits/208 ├── timesplit_no_lig_overlap_train209 ├── timesplit_no_lig_overlap_val210 └── timesplit_test211```212 213Split files should be plain text files, with one complex name per line. Each name must correspond to a subdirectory name in the data directory.214 215# Official OneScience Information216 217| Platform | OneScience main repository | Skills repository |218| --- | --- | --- |219| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |220| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |221 222# Citation and License223 224- The original DiffDock code is licensed under the MIT License. This repository retains source attribution and is organized for OneScience Hugging Face automated runtime scenarios.225- If you use DiffDock results in research, we recommend citing the original DiffDock and DiffDock-L papers, as well as relevant OneScience project information. Depending on the actual task, also add citations for datasets or tools such as PDBBind, MOAD, RDKit, OpenBabel, and GNINA.226 227```bibtex228@inproceedings{corso2023diffdock,229 title={DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking},230 author={Corso, Gabriele and St{\"a}rk, Hannes and Jing, Bowen and Barzilay, Regina and Jaakkola, Tommi},231 booktitle={International Conference on Learning Representations},232 year={2023}233}234```235 