Machine Learning and Chemometrics for Herbal Medicine Quality Control: Analytical Platforms, Performance Benchmarks and Barriers to Regulatory Adoption
Main Article Content
Abstract
The global herbal medicine market has expanded rapidly, exceeding an estimated USD 150 billion in valuation, yet quality control remains a critical challenge because of species misidentification, adulteration, and geographical origin fraud. Traditional quality control based on single-marker quantification fails to capture the chemical complexity of herbal matrices. Machine learning and chemometrics have emerged as transformative tools for pattern-based authentication and quality assessment. This review synthesised literature from SciSpace, PubMed, Google Scholar and ArXiv covering 2015–2025, following PRISMA-style selection criteria. Of 1,847 records identified, 450 underwent full-text review and 120 met the inclusion criteria, yielding a final corpus of 150 papers after citation chaining. Papers were classified by analytical platform (vibrational spectroscopy, chromatography–mass spectrometry, hyperspectral imaging, sensor fusion), by chemometric and machine learning method (classical multivariate, supervised learning, deep learning), and by application domain (authentication, adulteration, origin traceability, processing quality, bioactive quantification). Attenuated total reflectance Fourier transform infrared spectroscopy combined with support vector machines achieved 100% classification accuracy for 53 root and rhizome Chinese herbal species. Near-infrared spectroscopy coupled with kernel extreme learning machines gave 95.6% accuracy in American ginseng adulteration detection. Deep learning architectures, including one-dimensional convolutional neural networks and long short-term memory networks, outperformed classical chemometric methods in spectral feature extraction, with accuracies above 98% reported for hyperspectral imaging-based authentication. Data fusion combining laser-induced breakdown spectroscopy with Raman spectroscopy reached 93.4% accuracy through mid-level fusion, surpassing single-modality approaches. Metabolomics-driven quality marker discovery using backpropagation artificial neural networks yielded correlation coefficients above 0.99 for bioactivity prediction. Machine learning and chemometrics have therefore matured into robust tools that outperform traditional single-marker approaches. Critical gaps nonetheless persist in standardised benchmark datasets, model interpretability for regulatory acceptance, instrument-to-instrument transferability, and translation to portable field devices. Future pathways include explainable artificial intelligence, transfer learning for small-sample scenarios, multi-omics integration, blockchain-enabled supply chain traceability, and edge computing for real-time quality screening.
