J Comput Aided Mol Des. 2026 Aug 29;40(1):219. doi: 10.1007/s10822-026-00929-9.
ABSTRACT
Molecular docking is one of the most established methods in computational drug discovery, due to its balance of speed and accuracy. However, the accuracy of docking results depends on a number of different parameters, and systematic reference data for comparisons to more advanced methods for binding affinity prediction are still scarce. This study assesses the impact of key parameters on the accuracy of binding free energy estimates from docking, using nine benchmark systems with 278 high-affinity ligands. Using the Molecular Operating Environment (MOE), we evaluated combinations of three receptor structures (two crystal structures, one AlphaFold2 model), two force fields, two scoring functions, two receptor flexibility settings, and two statistical evaluation schemes. The performance of the docking approaches is measured based on the squared Pearson’s correlation coefficient (R²), the root mean square error (RMSE) with respect to the experimental binding affinities, as well as the mean signed error (MSE) and Kendall’s tau for individual targets and the full dataset. The results show that the scoring function and the protein structure are the most important factors for binding affinity accuracy in rigid docking with the MOE software. Amber10:EHT and MMFF94x force fields had the same average Rmean2 value, but Amber10:EHT had a lower average RMSEmean. AlphaFold2 protein models yielded lower binding affinity accuracy and higher errors compared to experimental crystal structures, although induced fit docking improved results. Using the original benchmark, we also compared several docking programs. DOCK6 and MOE performed best, with mean R² values of about 0.49 and 0.40, respectively. The remaining docking programs did not outperform a molecular weight regression baseline. For a subset of four targets (CDK2, JNK1, P38, TYK2) evaluated in previous work, the performance of the optimized DOCK6 and MOE protocols produced correlation coefficients similar to those reported for certain MM/PBSA, FMO, and Boltz2 implementations evaluated on the same target subset. This raises questions about potential dataset biases, the structural preparation, or the implementation of those methods. Docking therefore should be considered as an important and computationally inexpensive reference baseline for binding affinity prediction.
PMID:42667474 | DOI:10.1007/s10822-026-00929-9