Machine Learning Driven Discovery of Anticancer Phytochemicals from Medicinal Plants: A Computational Approach

Main Article Content

Humaira Shaheen, Rabia Shaheen Niazi, Sobia Niaz, Amber Sharif, Riffat Farooqui, Muhammad Salman, Ahmad Jawad Khan, Aafaq Ali, Abid Ejaz, Mian Jahan Zaib Rasheed

Abstract

Natural products remain an important source of chemically diverse molecules for anticancer drug discovery, but the large number of compounds reported in natural product databases makes systematic experimental screening difficult. Computational approaches can assist this process by identifying structural patterns associated with biological activity and prioritizing compounds for subsequent investigation. In this study, a machine-learning framework was developed for the computational prioritization of plant-derived compounds associated with cancer related bioactivity using data from NPASS v3.0. Plant-derived records were filtered against cancer associated biological activity, followed by molecular structure standardization, duplicate consolidation, activity label assignment, molecular descriptor calculation, fingerprint generation, and scaffold analysis. The final curated dataset contained 6,120 unique compounds representing 2,145 Bemis Murcko scaffolds, including 2,693 active and 3,427 inactive compounds. Molecular information was represented using 208 physicochemical and topological descriptors together with circular and structural fingerprints. Random forest, support vector machine, XGBoost, and deep neural network models were evaluated using an 80/10/10 scaffold aware train validation test design. Combining ECFP4 fingerprints with RDKit descriptors consistently improved predictive performance compared with individual representations. The deep neural network achieved the highest reported classification performance, with ROC-AUC of 0.883, PR-AUC of 0.857, accuracy of 0.817, F1-score of 0.782, and MCC of 0.627 on the held-out test set. Quantitative modelling also showed useful predictive performance for pIC , with the best DNN model obtaining a test-set  of 0.712, RMSE of 0.701, and MAE of 0.524. The scaffold-aware evaluation provides a more conservative estimate of molecular generalization than a conventional random split, although independent external validation remains necessary. Overall, the proposed framework provides a reproducible computational strategy for prioritizing plant-derived phytochemicals for subsequent anticancer investigation.

Article Details

Section
Articles