Hybrid CNN–Transformer Architecture for Multi-Organ Disease Classification

Main Article Content

Sachin Arun Thanekar, Ganesh Dagadu Puri, Vinodpuri Rampuri Gosavi, Yogesh Shepal, Bhausaheb R. Varpe, Shalini Wankhade

Abstract

Convolutional neural networks (CNNs) and Vision Transformers (ViTs) offer complementary strengths for medical image analysis: CNNs excel at capturing local texture and edge information through their inductive spatial bias, while transformers capture long-range dependencies through global self-attention but typically require more data to train effectively. Most existing diagnostic systems remain organ-specific, applying a purpose-built CNN, transformer, or hybrid model to a single body region — chest, brain, retina, breast, or skin — with little architectural sharing across organs, which duplicates engineering effort and prevents any given model from benefiting from cross-organ visual regularities such as shared texture statistics or common lesion morphology. This paper proposes a hybrid CNN-Transformer framework for multi-organ disease classification that uses a shared convolutional stem to extract local features, a transformer encoder to model global dependencies, and an organ-aware gated fusion module that combines both feature streams before routing them to organ-specific classification heads, allowing a single backbone to serve chest, brain, and retinal imaging simultaneously. We synthesize nineteen related studies spanning pure-CNN, pure-transformer, and hybrid architectures across seven organ systems (lung, breast, retina, brain, skin, thyroid, and middle ear), and identify the specific architectural gap this framework addresses: existing hybrid CNN-transformer designs are fused and validated within a single organ, leaving open whether a shared hybrid backbone with organ-aware fusion can generalize across organs without sacrificing the strong single-organ performance already demonstrated in the literature. Consistent with its status as a proposed framework, the Results and Discussion section positions expected performance against real, previously published benchmark figures for representative CNN, transformer, and hybrid models across organs, rather than reporting new experimental numbers, and this framing is stated explicitly throughout. The paper concludes with a three-phase empirical validation roadmap.

Article Details

Section
Articles