Deep Learning-Based Classification of Medicinal Plant Leaves: An Internal-Test Comparative Study of Betel, Curry Leaf, Aloe Vera, and Tulsi

Main Article Content

Mohan Kumara H K, K. Satyanarayan Reddy

Abstract

Accurate species-level identification of medicinal plant leaves is a foundational requirement for herbal quality assurance, precision cultivation, and downstream disease and quality-grading pipelines. This paper presents an automated leaf-image classification framework that categorizes four medicinally and economically significant Indian plant species — Betel (Piper betle), Curry Leaf (Murraya koenigii), Aloe Vera (Aloe vera (L.) Burm.f.; syn. Aloe barbadensis Mill.), and Tulsi (Ocimum tenuiflorum L.; syn. Ocimum sanctum L.) — into their respective classes. Two separate experimental tracks are evaluated on a common 1,179-image, single-source dataset with naturally varying backgrounds: (i) classical machine learning on handcrafted visual features, and (ii) end-to-end and transfer-learning convolutional neural networks; the two tracks are compared against one another rather than fused, and no hybrid or feature-fusion model is claimed. Under track (i), three classical classifiers — SVM, Random Forest, and ANN — were trained on a 1,834-dimensional handcrafted feature vector (color histograms, GLCM and LBP texture descriptors, HOG, and shape/morphology features) reduced via PCA fitted on the training split only, with hyperparameters selected by 5-fold cross-validation. The grid-searched SVM achieved the best held-out test performance among the classical classifiers at 56.5% accuracy (exact 95% binomial CI 48.9–63.9%; 57.9% macro-precision, 56.9% macro-recall, 56.6% macro-F1), with confusion-matrix analysis identifying Tulsi–Curry Leaf confusion as the dominant error source under cluttered natural backgrounds. Under track (ii), a CNN trained from scratch and four ImageNet-pretrained transfer-learning backbones (VGG-16, MobileNetV2, EfficientNet-B0, ResNet-50) — each fine-tuned by training a new classifier head on a frozen, ImageNet-pretrained convolutional base — were trained and evaluated on the identical 825/177/177 train/validation/test split on a GPU-enabled machine: the from-scratch CNN reached 73.4% test accuracy, VGG-16 reached 93.2%, MobileNetV2 and EfficientNet-B0 each reached 98.9%, and ResNet-50 reached 100% accuracy on the 177-image internal test set (exact 95% binomial CI 97.9–100%). These are internal-test results from a single random, image-level split of one public data source; they are not yet evidence of generalization to independently collected specimens, cameras, or field conditions, and independent, specimen-aware external validation is identified as a required next step before any deployment or field-readiness claim. The framework is intended as Stage 1 of a larger pipeline that would subsequently perform leaf-quality grading and disease-severity classification within the identified species.

Article Details

Section
Articles