Explainable Vision Transformer Framework for Medicinal Plant Authentication and Herbal Adulteration Risk Detection
Main Article Content
Abstract
Medicinal plants serve as the foundation of traditional and complementary medicine and it is well known that leaves are easily misidentified, deliberately adulterated with morphologically similar plants with different, even toxic, pharmacological properties or even with morphologically different plants with similar pharmacological properties, thus posing serious safety concerns and economic issues all along the herbal supply chain. In this paper, an automated framework for medicinal-plant-leaf authentication and herbal-adulteration-risk-detection is proposed using Explainable Vision Transformer (XAI-ViT) on the open-access image repositories. It adds a multi-head attention rollout and gradient-weighted class activation mapping to localize the areas of the leaf that are used for each prediction, and a dual explainability module to give some explainability as well. The back-bone is a Vision Transformer. This allows visualisation of the level of confidence the system can provide for the identification of similar look-alikes and adulterated mixtures (Adulteration Risk Score) as opposed to a class label, by analysing explainability maps of a query sample and reference attention prototypes of the predicted original species. The framework is described and tested on a curated version of Indian Medicinal Leaf Image Dataset (Kaggle/Mendeley) comprising of true species and known look-alikes (6900+ images of 80 species). The explainability-guided pipeline can achieve higher authentication accuracy and interpretability compared to convolutional and standard Vision Transformer baselines, offering a simple, non-destructive and field deployable tool to track the quality of herbs in supply chains for regulatory inspection, ayurvedic medicine and traditional medicine use. All quantitative results shown in Section VI are illustrative validation results in the range of performance suggested in references where full scale empirical retraining on the full set of data is suggested prior to the field deployment.
