PolyPharm is a deep learning-based framework for multi-target drug design, capable of generating molecules with potential activities against multiple targets.
/ ← Root directory
├── data ← Dataset, interaction scoring models, and multi-target scoring weights
│ ├── GSK3B+JNK3
│ └── ROR_gamma+DHODH
├── results ← Generated molecules from different methods on benchmark tasks
│ ├── gsk3β_jnk3
│ └── rorγt_dhodh
├── model ← Model code
├── score_modules ← Scoring utilities
│ ├── ESOL_Score
│ └── SA_Score
├── utils ← Utilities (multi-target scoring, data preprocessing, etc.)
├── environment.yml ← Conda environment
├── train_chembl_baseline.py ← Pre-training script
├── RL_generate.py ← Fine-tuning script
├── generate.py ← Molecule generation script
├── hyperparameter_search.py ← Fine-tuning parameter selection script
├── generate.py ← Molecule generation script
✅ Pre-generated molecules are provided for download (including QED, SA, Docking score, LogP, Weight) along with comparison methods (some from AIxFuse open-source data):
- GSK3β|JNK3 benchmark task:
results/gsk3β_jnk3/POLYGEN.csv - RORγt|DHODH benchmark task:
results/rorγt_dhodh/POLYGEN.csv
💡 To train from scratch, follow the steps below.
git clone https://github.com/Yozu-Roo/POLYGEN.git- Dataset & other tools (~1.69GB) download via Google Drive
- Or download via Baidu Drive
- Or contact via email: niuziru@stu.xmu.edu.cn
After download, extract and place the data/ folder at the root directory.
Recommended Python 3.8:
conda env create -f environment.yml
conda activate polypharmpython train_chembl_baseline.py- For multiple GPUs, adjust
CUDA_VISIBLE_DEVICESin the script - Model weights are saved in
pretrain_output/during training
Splitting the fine-tuning dataset:
python utils/split_finetune_dataset.py \
--ligands_set1 ./data/GSK3B+JNK3/GSK3B.csv \
--ligands_set2 ./data/GSK3B+JNK3/JNK3.csv \
--output_dir ./data/GSK3B+JNK3/ \
--train_ratio 0.7 \
--val_ratio 0.1 \
--test_ratio 0.2python utils/split_finetune_dataset.py \
--ligands_set1 ./data/ROR_gamma+DHODH/ROR_gamma.csv \
--ligands_set2 ./data/ROR_gamma+DHODH/DHODH.csv \
--output_dir ./data/ROR_gamma+DHODH/ \
--train_ratio 0.7 \
--val_ratio 0.1 \
--test_ratio 0.2The train.csv, val.csv, and test.csv—will be saved in the output_dir.
Fine-tune on GSK3β|JNK3 benchmark task:
python RL_generate.py \
--target_name GSK3B JNK3 \
--output_dir ./finetune_output_GJ \
--data_path ./data/GSK3B+JNK3/train.csv \
--model_path ./pretrain_output/fold0_epoch32.pth \
--tokenizer_path ./pretrain_output/tokenizer.pkl \
--n_mol 10000 \
--device cuda \
--batch_size 512 \
--seed 42 \
--threshold 0.60 \
--n_epochs 20 \
--optimize_n_epochs 5 \
--save_frequency 10 \
--save_payloads \
--keep_top_ratio 0.5Fine-tune on RORγt|DHODH benchmark task:
python RL_generate.py \
--target_name ROR_gamma DHODH \
--data_path ./data/ROR_gamma+DHODH/train.csv \
--output_dir ./finetune_output_RD \
--model_path ./pretrain_output/fold0_epoch32.pth \
--tokenizer_path ./pretrain_output/tokenizer.pkl \
--n_mol 10000 \
--device cuda \
--batch_size 512 \
--seed 42 \
--threshold 0.60 \
--n_epochs 15 \
--optimize_n_epochs 5 \
--save_frequency 10 \
--save_payloads \
--keep_top_ratio 0.5--tokenizer_pathis your pre-trained model path--thresholdis the threshold for screening elite molecules--n_epochsis the fine-tuning epochs--optimize_n_epochsis the optimization epochs--n_molis the number of molecules sampled each epoch- Generated results and model weights are saved in
finetune_output_*/ - Multi-GPU users may modify
CUDA_VISIBLE_DEVICES
Note
⚙️ Hyperparameter selection
Hyperparameters were selected separately for the two benchmark tasks. We evaluated the number of fine-tuning epochs (n_epochs) over the range of 5–25 with a step size of 5, while fixing threshold at 0.60. We then evaluated the elite-molecule screening threshold (threshold) over the range of 0.50–0.70 with a step size of 0.05, while fixing n_epochs at 20. The remaining fine-tuning settings were kept unchanged.
The entire hyperparameter search can be performed automatically using:
python hyperparameter_search.py \
--target_name ROR_gamma DHODH \
--data_path ./data/ROR_gamma+DHODH/train.csv \
--val_data_path ./data/ROR_gamma+DHODH/val.csv \
--model_path ./pretrain_output/fold0_epoch32.pth \
--tokenizer_path ./pretrain_output/tokenizer.pkl \
--output_prefix ./hyperparam_RD \
--n_mol 30000 \
--generate_n_mol 10000 \
--device cuda \
--batch_size 512 \
--seed 42 \
--optimize_n_epochs 5 \
--save_frequency 10 \
--keep_top_ratio 0.5For the GSK3β|JNK3 benchmark task:
python hyperparameter_search.py \
--target_name GSK3B JNK3 \
--data_path ./data/GSK3B+JNK3/train.csv \
--val_data_path ./data/GSK3B+JNK3/val.csv \
--model_path ./pretrain_output/fold0_epoch32.pth \
--tokenizer_path ./pretrain_output/tokenizer.pkl \
--output_prefix ./hyperparam_GJ \
--n_mol 10000 \
--generate_n_mol 10000 \
--device cuda \
--batch_size 512 \
--seed 42 \
--optimize_n_epochs 5 \
--save_frequency 10 \
--keep_top_ratio 0.5After the script finishes, CSV files containing the generated molecules for different parameter settings will be produced in the following directories:
hyperparam_**_threshold_**_generation/ and hyperparam_**_epoch_**_generation/.
Next, perform batch docking on the generated molecules using AutoDock Vina and calculate their QED and SA properties. Finally, calculate the USR docking scores and SR values according to the evaluation metrics described in the manuscript. Compare these results and select the parameter combination with the highest SR and USR docking scores as the final hyperparameter setting.
python generate.py \
--target_name ROR_gamma DHODH \
--output_dir ./generate_output_RD \
--save_file RD_gen.csv \
--model_path ./finetune_output_RD/epoch_15_finetuned_model.pth \
--tokenizer_path ./pretrain_output/tokenizer.pkl \
--init_smi_path ./data/ROR_gamma+DHODH/test.csv \
--n_mol 10000 \
--device cuda \
--filter \
--batch_size 512 \
--seed 42--target_nameis the benchmark task and can be replaced withGSK3B JNK3--init_smi_pathis the file path for the test set.--model_pathis your fine-tuned model path--n_molis the number of generated molecules- Generated molecules are saved in
generate_output_*/
- Ensure all paths are correct to avoid file-not-found errors
- GPU significantly speeds up training and generation
- Docking tool: AutoDock Vina or using Vina-GPU speeds up
- Retrosynthesis tool: AiZynthFinder