Training Seed Variance Exceeds Architectural Differences in Lightweight Diabetic Retinopathy Grading

Main Article Content

Himansu Sekhar Rout, Saravana Kumar S

Abstract

Architectural choices in automated diabetic retinopathy grading are routinely justified by a single training run on a single benchmark. This paper asks whether the differences such comparisons report are large enough to be interpretable. Four variants of a compact grading network, differing only in how encoder features are aggregated, together with an EfficientNet-B1 baseline, were each trained five times from different random seeds under otherwise identical conditions on APTOS 2019, giving twenty-five runs, and every run was evaluated both on a frozen APTOS test partition and, with weights unchanged, on the independent IDRiD dataset. Between-architecture differences are smaller than the variation between repeated runs of the same architecture. The spread between architecture means is 0.012 quadratic weighted kappa (QWK) on APTOS and 0.013 on IDRiD, against median seed-to-seed standard deviations of 0.010 and 0.027 respectively; on the external dataset the noise is twice the signal. Four of the five architectures achieve the best score on at least one seed on each dataset, and most span the full range from first to last. A single-seed experiment conducted first appeared to show the ranking reversing between the two datasets, with an architecture-by-dataset interaction of +0.052 QWK (p = 0.012); across five seeds that interaction has mean −0.006 and standard deviation 0.039, and no model pair shows a stable interaction. Findings that do survive replication are reported: all models compress predictions toward the middle of the ordinal scale under transfer, mild disease collapses from an F1 of 0.46–0.60 to 0.14–0.18, and ImageNet pretraining separates from scratch training by roughly four times the observed seed variance. The practical conclusion is that architectural rankings in this regime require repeated training to be meaningful, and that single-run comparisons, including those in recent work, should be read as provisional.

Article Details

Section
Articles