Explainable Vision Transformer Framework for Early Multi-Disease Detection from Multimodal Medical Images

Main Article Content

Shilpa Nikhil Bhosale, Supriya Sabale, Yogesh Shepal, Aarti S. Gaikwad, Bhausaheb R. Varpe, Chandrakant D. Kokane

Abstract

Chest radiography (CXR) and computed tomography (CT) remain the two most accessible imaging modalities for screening respiratory and thoracic disease, yet the deep learning systems built around them are typically narrow in three respects: they consume a single modality, target a single disease category and offer clinicians little more than a probability score with no visual justification. This combination limits both early multi-disease detection and clinical trust, since radiologists cannot verify why a model reached a given conclusion. This paper proposes an Explainable Vision Transformer (ViT) framework that fuses paired chest X-ray and CT representations through a cross-modal co-attention mechanism, performs multi-label classification across several thoracic conditions simultaneously and produces per-prediction visual explanations by combining attention-rollout and Grad-CAM-style localization adapted to transformer encoders. We synthesize twenty related studies spanning CNN-based, ViT-based and multimodal fusion literature, identify the specific architectural and explainability gaps that motivate this work and detail a complete methodology, training protocol and evaluation plan. Because this paper is presented as a proposed framework and literature-grounded roadmap rather than a completed empirical study, the Results and Discussion section positions the framework against real, previously published benchmark figures for comparable multimodal and explainable chest-imaging models, rather than reporting new experimental numbers. This distinction is stated explicitly so that readers can correctly interpret the performance context. The paper concludes with a discussion of expected trade-offs, limitations of literature-based benchmarking and a concrete plan for the empirical validation phase.

Article Details

Section
Articles