An Integrated Molecular Docking and Machine-Learning Framework for Prioritising Natural Anticancer Leads from South Indian Medicinal Plants against Multiple Validated Cancer Targets
Main Article Content
Abstract
Background: The chemical diversity of South Indian medicinal plants offers a large, under-exploited reservoir of potential anticancer scaffolds, yet the sheer number of phytochemical-target combinations makes purely experimental screening inefficient. We developed an integrated in-silico framework coupling structure-based molecular docking with supervised machine learning (ML) and predicted absorption, distribution, metabolism, excretion and toxicity (ADMET) profiling to rationally prioritise natural leads. Methods: A curated library of 75 phytochemicals from 15 South Indian medicinal plants was assembled and profiled against 10 experimentally validated cancer targets (EGFR, VEGFR2, PIK3CA, AKT1, CDK2, BCL2, MDM2, PARP1, HDAC2 and TOP2A). Molecular descriptors and ADMET properties were computed with RDKit. Binding affinities were generated for all 750 compound-target pairs. Four ML classifiers (random forest, gradient boosting, support-vector machine and logistic regression) were trained under stratified five-fold cross-validation (seed 42) to predict potent binders from 13 physicochemical descriptors. A consensus score combining normalised docking (0.45), calibrated ML probability (0.40) and inverse ADMET risk (0.15) ranked integrated hits. Results: Docking affinities spanned -7.21 to -12.50 kcal/mol (mean -9.74); 275 of 750 pairs reached <=-10 kcal/mol. The random forest was the best classifier (ROC-AUC 0.912, accuracy 0.853, MCC 0.692). Twenty-five compound-target pairs were classified as high-priority. Top-ranked leads included acetyl-beta-boswellic acid (Boswellia serrata)-EGFR, chebulagic acid (Terminalia chebula)-BCL2 and corilagin-MDM2. Sixty-four of 75 compounds satisfied Lipinski criteria (<=1 violation). Conclusions: The integrated docking-ML-ADMET pipeline reproducibly converged on ellagitannin, triterpene and boswellic-acid chemotypes as priority anticancer leads, providing a transparent, computationally generated hypothesis set for subsequent experimental validation.
