Digital Twin-Driven Explainable Generative AI Framework for Personalized Disease Progression Prediction and Precision Healthcare

Main Article Content

Sachin Arun Thanekar, Ganesh Dagadu Puri, Deepali Mahesh Gohil, Dewanand Atmaram Meshram, Baliram Sambhaji Gayal, Sunita Nandgave

Abstract

Chronic disease management increasingly demands prediction tools that go beyond a single point estimate of risk and instead offer a personalized, time-evolving and uncertainty-aware picture of how an individual patient's condition is likely to unfold. This paper proposes a Digital Twin-Driven Explainable Generative AI (DT-XGAI) framework for personalized disease progression prediction, in which each patient's baseline clinical state is used to instantiate a probabilistic "digital twin" — a generative model capable of simulating a distribution of plausible future disease trajectories rather than a single deterministic forecast — coupled with a dual explainability layer combining feature-attribution and model-agnostic permutation-based rationale.
We instantiate and rigorously evaluate a proof-of-concept version of this framework on the open scikit-learn Diabetes Progression dataset (Efron et al., 2004): 442 patients described by ten baseline physiological and serum biomarker variables, with a genuine quantitative one-year disease-progression outcome. A Gaussian-Process-based generative digital twin core is compared against Linear Regression, Random Forest and Gradient Boosting baselines under five-fold cross-validation. The generative digital twin achieved the best overall cross-validated performance (R² = 0.486 ± 0.061, MAE = 43.88 ± 1.97), modestly exceeding the Linear Regression baseline (R² = 0.478) and clearly outperforming Random Forest (R² = 0.442) and Gradient Boosting (R² = 0.427), while additionally providing calibrated, patient-specific predictive uncertainty that none of the deterministic baselines can produce.
On a held-out test set, the digital twin achieved R² = 0.496 and MAE = 41.15 and posterior-sample trajectory simulation for representative patients is shown to produce personalized 95% confidence bands whose width varies meaningfully across individuals. SHAP-based attribution on a Gradient Boosting surrogate and permutation-importance analysis on the generative model independently converge on body mass index (BMI) and a serum lipid biomarker (s5) as the two dominant drivers of predicted progression, consistent with the original clinical findings on this cohort. These results provide honest, reproducible, small-scale evidence that generative, uncertainty-quantifying digital twin models can match or exceed deterministic point-prediction baselines while offering substantially richer, more clinically actionable output. We discuss the current proof-of-concept's limitations, outline a roadmap toward multimodal, longitudinal, credentialed clinical validation and argue that uncertainty-aware, explainable digital twins — not single-number risk scores — represent a more defensible foundation for precision healthcare decision support.

Article Details

Section
Articles