Stability-Guided Genetic Feature Selection Produces a Compact Cardiovascular Disease Classifier with Preserved Discrimination
Main Article Content
Abstract
Feature selection can reduce measurement burden in cardiovascular disease classification, but single-run evolutionary searches may be unstable and may overstate gains when feature selection and evaluation are not separated. We analyzed a public cardiovascular dataset containing 70,000 examination records. After removing duplicate profiles and applying prespecified physiological plausibility rules, 68,586 records were retained. Six clinically interpretable variables were engineered, giving 16 candidate predictors. A stability-guided, cost-aware genetic algorithm was run five times on training data, and features selected in at least two runs formed an eight-variable consensus panel. Logistic regression, Gaussian naive Bayes, random forest, and LightGBM were evaluated before and after selection; a validation-trained stacked ensemble was assessed secondarily. Final estimates were obtained from an untouched test set of 10,288 records with bootstrap confidence intervals, calibration, decision-curve analysis, and subgroup auditing. The consensus panel contained age, weight, systolic and diastolic blood pressure, cholesterol, physical activity, pulse pressure, and mean arterial pressure, reducing the feature set by 50%. Full-feature LightGBM achieved ROC-AUC 0.8027 (95% CI 0.7948-0.8109), while GA-LightGBM achieved 0.8003 (95% CI 0.7923-0.8088), accuracy 0.7341, and Brier score 0.1811, with an 18.6% reduction in measured training time. The AUC decrease was small but statistically detectable. The stacked model achieved AUC 0.7998 and did not improve on GA-LightGBM. Stability-guided consensus therefore produced a substantially smaller, clinically coherent panel while preserving nearly all discrimination, supporting parsimony with independent evaluation, calibration, and external validation.
