MEFS: Multistage Ensemble Feature Selection Framework for Diabetes Risk Prediction Using Relevance, Separability, and Feature Diversity

Main Article Content

Ashvini Sandip Patil, Rahul Raghvendra Joshi, Milind Gayakwad

Abstract

Diabetes mellitus is among the most common metabolic diseases and requires early and accurate diagnosis in order to avoid any additional deterioration and complications.  The performance of machine learning algorithms is promising in diabetes prediction, but their efficiency is usually hindered by redundant features, class imbalance, and poor interpretability of the developed models. Common feature selection methods usually employ only a relevance measure and neglect to incorporate nonlinear dependency, separability, and diversity, leading to poor predictive efficiency and high computational cost. To overcome these problems, we propose Multistage Ensemble Feature Selection (MEFS) framework which combines Nonlinear Dynamic Conditional Relevance Feature Selection (NDCRFS), Hellinger Distance (HD), and Feature Hamming Diversity (FHD) to select an optimal set of informative features for diabetes predictions. Firstly, the data is preprocessed using Feature-Specific Mean Variance Normalization (FSMVN) and Adaptive Synthetic Sampling (ADASYN). Then, the proposed feature selection method utilizes nonlinear analysis, distribution comparison, and redundancy elimination using an ensemble selection process. The proposed feature selection framework subsequently combines nonlinear relevance analysis, distribution-based discrimination, and redundancy reduction through an ensemble ranking strategy.  Performance evaluation of the feature subset is conducted using SVM, LR, RF, XGBoost, and MLP classifiers with ten-fold cross-validation on the Pima Indians Diabetes dataset. From the experimental analysis, it can be seen that the MSFS helps in reducing the size of the original feature space by 37.5%, utilizes lesser memory with an approximation of 37% memory savings, and helps in saving computational training time without affecting prediction accuracy to any significant extent. The statistical verification process conducted through the Friedman test and Wilcoxon signed rank test proves that the suggested methodology performs as well as existing methods of feature selection without deteriorating performance.

Article Details

Section
Articles