HybridDermNet: A CNN–Transformer Framework with Latent Bottleneck Learning for Robust Skin Cancer Classification
Main Article Content
Abstract
The classification of skin lesions from dermoscopic and clinical images remains challenging due to visually similar lesion categories, variability in image acquisition, class imbalance, and limited cross-dataset generalisation, which can undermine the reliability of diagnostic results. Current CNN-based approaches are good at local texture, pigmentation, and border information, while transformer models are good at global structure and long-range dependencies. Most hybrid methods, however, use simple concatenation or additive fusion, or rely on multi-backbone architectures that are very computation-intensive, produce redundant features, and risk overfitting. In this work, a CNN–Transformer-based framework with latent bottleneck learning for robust multi-class skin cancer classification is proposed. The algorithm locally and globally extracts features in parallel, calculates bidirectional attention between features, adaptively fuses features, encodes features in a compact latent space, optionally reconstructs features, and predicts with confidence. A composite objective combines classification, reconstruction, cross-feature consistency, latent regularisation, and transformation-consistency loss. HybridDermNet achieved 96.18% accuracy, 94.86% macro-F1, and 97.92% AUROC on HAM10000, and 95.42% accuracy, 93.92% macro-F1, and 97.46% AUROC on ISIC 2018. When transferring between datasets, it achieved 87.24% and 87.61% macro-F1, respectively, and, in external validation on PAD-UFES-20, achieved 80.67% macro-F1. The impact of attention-guided fusion and latent compression was verified through ablation, calibration, explainability, efficiency, and statistical analysis. The framework offers compact, interpretable and generalizable representations for reliable dermatology decision support in heterogeneous imaging conditions.
