Diagnostic Accuracy of Artificial Intelligence in Classifying Oral Ulcerative Lesions in Clinical Practice: A Comparative Clinical Study
Main Article Content
Abstract
Background: Oral ulcerative lesions encompass a diverse range of conditions with overlapping clinical manifestations, making accurate differential diagnosis challenging. Artificial intelligence (AI)-based image analysis has emerged as a promising approach for supporting the recognition and classification of oral lesions. The present study evaluated the diagnostic performance of an AI model for classifying oral ulcerative lesions using clinical images.
Aim: To assess the diagnostic accuracy of an AI model in differentiating major categories of oral ulcerative lesions and to compare its performance with conventional clinical assessment.
Materials and Methods: A total of 150 patients with clinically diagnosed oral ulcers were included in the study. Clinical photographic images were analyzed by the AI model and classified into six diagnostic categories: aphthous, traumatic, herpetic, fungal, tuberculous, and ulcerative oral squamous cell carcinoma (OSCC). The final clinical and/or histopathological diagnosis was considered the reference standard. Diagnostic performance was assessed using accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1-score, receiver operating characteristic (ROC) analysis, and Cohen’s kappa (κ). The diagnostic performance of the AI model was also compared with conventional clinical assessment.
Results: The AI model demonstrated strong overall diagnostic performance in the classification of oral ulcerative lesions. It achieved an overall accuracy of 92.7%, sensitivity of 92.3%, specificity of 97.9%, and an ROC-AUC of 0.965. Agreement with the reference diagnosis was excellent, with a κ of 0.904. Across the individual diagnostic categories, sensitivity and specificity remained consistently high, with the strongest classification performance observed for herpetic and tuberculous ulcers. The AI model significantly outperformed conventional clinical assessment in terms of diagnostic accuracy, sensitivity, specificity, F1-score, and discriminatory ability.
Conclusion: The findings indicate that AI-based analysis of clinical images can provide accurate and consistent classification of oral ulcerative lesions and may serve as a valuable adjunct to conventional clinical assessment. Its potential to support differential diagnosis and early clinical decision-making warrants further evaluation through larger, prospective, multicenter, and externally validated studies.
