Skip to content

Repository files navigation

🎯 PolyPharm

PolyPharm is a deep learning-based framework for multi-target drug design, capable of generating molecules with potential activities against multiple targets.

📁 Directory Structure & Key Files

/                         ← Root directory
├── data                  ← Dataset, interaction scoring models, and multi-target scoring weights
│   ├── GSK3B+JNK3 
│   └── ROR_gamma+DHODH
├── results               ← Generated molecules from different methods on benchmark tasks
│   ├── gsk3β_jnk3 
│   └── rorγt_dhodh
├── model                 ← Model code
├── score_modules         ← Scoring utilities
│   ├── ESOL_Score
│   └── SA_Score
├── utils                     ← Utilities (multi-target scoring, data preprocessing, etc.)
├── environment.yml           ← Conda environment
├── train_chembl_baseline.py  ← Pre-training script
├── RL_generate.py            ← Fine-tuning script
├── generate.py               ← Molecule generation script
├── hyperparameter_search.py  ← Fine-tuning parameter selection script
├── generate.py               ← Molecule generation script

📊 Results

✅ Pre-generated molecules are provided for download (including QED, SA, Docking score, LogP, Weight) along with comparison methods (some from AIxFuse open-source data):

  • GSK3β|JNK3 benchmark task: results/gsk3β_jnk3/POLYGEN.csv
  • RORγt|DHODH benchmark task: results/rorγt_dhodh/POLYGEN.csv

💡 To train from scratch, follow the steps below.

🚀 Quick Start

📥 Clone Repository & Download Dataset

git clone https://github.com/Yozu-Roo/POLYGEN.git

After download, extract and place the data/ folder at the root directory.


📦 Install Conda Environment

Recommended Python 3.8:

conda env create -f environment.yml
conda activate polypharm

🧠 Pre-training

python train_chembl_baseline.py
  • For multiple GPUs, adjust CUDA_VISIBLE_DEVICES in the script
  • Model weights are saved in pretrain_output/ during training

🔧 Fine-tuning

Splitting the fine-tuning dataset:

python utils/split_finetune_dataset.py \
    --ligands_set1 ./data/GSK3B+JNK3/GSK3B.csv \
    --ligands_set2 ./data/GSK3B+JNK3/JNK3.csv \
    --output_dir ./data/GSK3B+JNK3/ \
    --train_ratio 0.7 \
    --val_ratio 0.1 \
    --test_ratio 0.2
python utils/split_finetune_dataset.py \
    --ligands_set1 ./data/ROR_gamma+DHODH/ROR_gamma.csv \
    --ligands_set2 ./data/ROR_gamma+DHODH/DHODH.csv \
    --output_dir ./data/ROR_gamma+DHODH/ \
    --train_ratio 0.7 \
    --val_ratio 0.1 \
    --test_ratio 0.2

The train.csv, val.csv, and test.csv—will be saved in the output_dir.

Fine-tune on GSK3β|JNK3 benchmark task:

python RL_generate.py \
    --target_name GSK3B JNK3 \
    --output_dir ./finetune_output_GJ \
    --data_path ./data/GSK3B+JNK3/train.csv \
    --model_path ./pretrain_output/fold0_epoch32.pth \
    --tokenizer_path ./pretrain_output/tokenizer.pkl \
    --n_mol 10000 \
    --device cuda \
    --batch_size 512 \
    --seed 42 \
    --threshold 0.60 \
    --n_epochs 20 \
    --optimize_n_epochs 5 \
    --save_frequency 10 \
    --save_payloads \
    --keep_top_ratio 0.5

Fine-tune on RORγt|DHODH benchmark task:

python RL_generate.py \
    --target_name ROR_gamma DHODH \
    --data_path ./data/ROR_gamma+DHODH/train.csv \
    --output_dir ./finetune_output_RD \
    --model_path ./pretrain_output/fold0_epoch32.pth \
    --tokenizer_path ./pretrain_output/tokenizer.pkl \
    --n_mol 10000 \
    --device cuda \
    --batch_size 512 \
    --seed 42 \
    --threshold 0.60 \
    --n_epochs 15 \
    --optimize_n_epochs 5 \
    --save_frequency 10 \
    --save_payloads \
    --keep_top_ratio 0.5
  • --tokenizer_path is your pre-trained model path
  • --threshold is the threshold for screening elite molecules
  • --n_epochs is the fine-tuning epochs
  • --optimize_n_epochs is the optimization epochs
  • --n_mol is the number of molecules sampled each epoch
  • Generated results and model weights are saved in finetune_output_*/
  • Multi-GPU users may modify CUDA_VISIBLE_DEVICES

Note

⚙️ Hyperparameter selection

Hyperparameters were selected separately for the two benchmark tasks. We evaluated the number of fine-tuning epochs (n_epochs) over the range of 5–25 with a step size of 5, while fixing threshold at 0.60. We then evaluated the elite-molecule screening threshold (threshold) over the range of 0.50–0.70 with a step size of 0.05, while fixing n_epochs at 20. The remaining fine-tuning settings were kept unchanged.

The entire hyperparameter search can be performed automatically using:

python hyperparameter_search.py \
    --target_name ROR_gamma DHODH \
    --data_path ./data/ROR_gamma+DHODH/train.csv \
    --val_data_path ./data/ROR_gamma+DHODH/val.csv \
    --model_path ./pretrain_output/fold0_epoch32.pth \
    --tokenizer_path ./pretrain_output/tokenizer.pkl \
    --output_prefix ./hyperparam_RD \
    --n_mol 30000 \
    --generate_n_mol 10000 \
    --device cuda \
    --batch_size 512 \
    --seed 42 \
    --optimize_n_epochs 5 \
    --save_frequency 10 \
    --keep_top_ratio 0.5

For the GSK3β|JNK3 benchmark task:

python hyperparameter_search.py \
    --target_name GSK3B JNK3 \
    --data_path ./data/GSK3B+JNK3/train.csv \
    --val_data_path ./data/GSK3B+JNK3/val.csv \
    --model_path ./pretrain_output/fold0_epoch32.pth \
    --tokenizer_path ./pretrain_output/tokenizer.pkl \
    --output_prefix ./hyperparam_GJ \
    --n_mol 10000 \
    --generate_n_mol 10000 \
    --device cuda \
    --batch_size 512 \
    --seed 42 \
    --optimize_n_epochs 5 \
    --save_frequency 10 \
    --keep_top_ratio 0.5

After the script finishes, CSV files containing the generated molecules for different parameter settings will be produced in the following directories: hyperparam_**_threshold_**_generation/ and hyperparam_**_epoch_**_generation/.

Next, perform batch docking on the generated molecules using AutoDock Vina and calculate their QED and SA properties. Finally, calculate the USR docking scores and SR values according to the evaluation metrics described in the manuscript. Compare these results and select the parameter combination with the highest SR and USR docking scores as the final hyperparameter setting.


🧬 Molecule Generation

python generate.py \
    --target_name ROR_gamma DHODH \  
    --output_dir ./generate_output_RD \
    --save_file RD_gen.csv \
    --model_path ./finetune_output_RD/epoch_15_finetuned_model.pth \  
    --tokenizer_path ./pretrain_output/tokenizer.pkl \
    --init_smi_path ./data/ROR_gamma+DHODH/test.csv \
    --n_mol 10000 \
    --device cuda \
    --filter \
    --batch_size 512 \
    --seed 42
  • --target_name is the benchmark task and can be replaced with GSK3B JNK3
  • --init_smi_path is the file path for the test set.
  • --model_path is your fine-tuned model path
  • --n_mol is the number of generated molecules
  • Generated molecules are saved in generate_output_*/

🌟 Tips

  • Ensure all paths are correct to avoid file-not-found errors
  • GPU significantly speeds up training and generation
  • Docking tool: AutoDock Vina or using Vina-GPU speeds up
  • Retrosynthesis tool: AiZynthFinder

About

PolyPharm: A deep learning-based framework for multi-target drug design

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages