Natural Product-Likeness Scores Revisited: Machine Learning Models for Predicting Drug-Like Properties from Phytochemical Databases

Main Article Content

Nirali Patel, Neha P Singh, Manvinder Brar, K.V. Navin Raja, Satish N. Pawar, Dilnavoz Jalilova, Vinitha M, Seethaladevi S

Abstract

Natural products remain among the richest sources of drug leads, yet the cheminformatic tools used to prioritise molecules were largely developed with synthetic compounds in mind. Natural products occupy a distinct region of chemical space—characteristically higher in molecular weight, oxygen content, fraction of sp³ carbons and stereocentres, and lower in nitrogen, lipophilicity and aromaticity than synthetic drugs—and roughly one in five violate Lipinski's “Rule of Five” while remaining bioactive. This mismatch motivates a fresh look at the two families of scores that guide natural-product-based discovery: drug-likeness metrics (the Rule of Five, Veber, Ghose, Egan, Muegge and the quantitative estimate of drug-likeness, QED) and the natural product-likeness (NP-likeness) score of Ertl and colleagues, a Bayesian fragment-based measure that cleanly separates natural products from synthetic molecules. This review revisits these scores in the machine-learning era.
We summarise the classical rules and their limitations for natural products; describe the NP-likeness score and its open-source and neural-network reimplementations; quantify the systematic physicochemical differences between natural-product and synthetic chemical space; survey the major phytochemical databases (COCONUT, LOTUS, NPASS, KNApSAcK) that now supply training data at scale; and examine how molecular representations (physicochemical descriptors, fingerprints such as ECFP/Morgan, and graph- and SMILES-based encodings) feed machine-learning models—random forests, gradient boosting, deep and graph neural networks—that predict drug-likeness, ADMET and bioactivity. We conclude that data-driven, natural-product-aware models are superseding rigid rule-based filters, but that their reliability depends on data quality, an honestly defined applicability domain, interpretability, and resistance to the historical bias toward synthetic, Rule-of-Five-compliant chemistry.

Article Details

Section
Articles