Self-Supervised Medical Image Analysis Using Foundation Models for Rare Disease Diagnosis
Main Article Content
Abstract
Rare diseases are individually uncommon but collectively affect a substantial share of patients, and their diagnosis from medical images is fundamentally constrained by the scarcity of labeled examples available for any single condition. Conventional supervised deep learning, which requires hundreds to thousands of labeled cases per class, is poorly matched to this setting and tends to overfit or fail to generalize when trained on the handful of cases a single institution may hold for a given rare condition. This paper proposes a self-supervised foundation model framework for rare-disease image analysis that pretrains a Vision Transformer encoder on large volumes of unlabeled medical images, spanning both common and rare conditions and multiple imaging domains, using a combination of teacher-student self-distillation and masked image modeling, and then adapts this encoder to specific rare-disease tasks using only a handful of labeled examples per class through linear probing, prompt-tuning, or lightweight adapters, together with an embedding-based uncertainty and case-retrieval mechanism to support clinician review. We synthesize twenty-two related studies spanning general-purpose self-supervised vision foundation models, medical-imaging-specific self-supervised encoders, and rare-disease-focused foundation models across fundus, chest imaging, and histopathology domains, and identify the specific gap this framework addresses: the near-absence of frameworks that combine domain-general self-supervised pretraining, explicit few-shot adaptation, and calibrated uncertainty estimation within a single pipeline purpose-built for rare-disease diagnosis. Consistent with its status as a proposed framework rather than a completed empirical study, the Results and Discussion section positions expected performance against real, previously published benchmark figures for comparable self-supervised and few-shot medical imaging systems, and this framing is stated explicitly throughout. The paper concludes with a three-phase empirical validation roadmap.
