Automated Detection, Quantification and Grading of Tuberculosis Bacilli in Ziehl-Neelsen Sputum Smear Microscopy: A Structured Review of Methods, Datasets, Multi-Centre Learning and Robustness

Main Article Content

Vishakha Yadav, Thippeswamy G

Abstract

Sputum smear microscopy remains the frontline diagnostic test for pulmonary tuberculosis in high-burden settings, yet it is slow, demands a trained microscopist and reports limited sensitivity. Deep learning has been applied intensively to this task, but the literature has grown without a shared benchmark, a shared metric definition or a shared statement of the clinical target. This review organises the 2022-2026 literature into a six-category taxonomy spanning acquisition and image formation, detection and segmentation, quantification and bacillary load grading, learning under limited supervision, multi-centre privacy-preserving learning, and robustness. Primary studies, seven reviews and every reachable public dataset are compared in eight tables. Three findings emerge. Reported performance is inflated by task substitution, since patch classification studies report accuracies near ninety-nine percent whereas detection studies report mean average precision near fifty percent when averaged across localisation thresholds. Quantification is largely unsolved, because clumped bacilli are annotated away rather than resolved and only one study automates the five-point World Health Organization load scale. Multi-centre and robustness evidence is almost absent, with one of one hundred and fifty-two artificial intelligence tuberculosis studies reporting a domain-shift analysis and no published federated learning study on sputum smear microscopy. We formalise the evaluation gap, propose a foreground-density heterogeneity measure for federated smear analysis, and set out a research agenda centred on grading-aware, multi-centre and reliability-audited models.

Article Details

Section
Articles