AI/ML Based Prediction of Experimentally Derived Bioactivity in Natural Products: A Systematic Review of Models, Validation Practices, and Performance Reporting
Main Article Content
Abstract
Researchers increasingly use AI/ML to predict bioactivity of natural products from molecular structure, but datasets, validation designs, and reporting practices vary widely across studies, making it hard to judge how trustworthy reported performance is? We conducted a systematic review of scholarly databases from 2015 to 2026 to identify AI/ML studies that predict experimentally measured bioactivity of natural products. Two reviewers screened records, extracted data, and applied an AI focused risk of bias framework covering dataset size, data leakage, validation design, applicability domain assessment, and reproducibility. Model performance was summarised using median AUC and related metrics from held out, external, or prospective validation, grouped by validation tier, dataset size, and model family. We included 16 studies (52 predictive models) among 2,705 records, spanning diverse natural product sources and pharmacological targets most models used classical machine learning. while deep learning and graph based methods were mainly used for larger or more complex datasets. Feature selection performed before data splitting and the absence of external validation were judged high risk in 56.3% studies. Models trained on small datasets with fewer than 100 compounds showed the highest median best AUC (0.97). whereas studies in higher validation tiers (Tier 4-5) often reported very high median best AUC (0.98-0.999) despite limited robust external testing, suggesting performance inflation in some settings. Overall, our findings indicate that improving validation design, explicitly analysing applicability domains, and routinely sharing code and data are likely to be more important for reliable natural product bioactivity prediction than further escalation in model complexity.
