Causal-GAT-LLM: A Hybrid Framework for Explainable Heart Disease Prediction using LiNGAM and Language Models
Main Article Content
Abstract
Heart disease remains a leading cause of mortality worldwide, necessitating accurate and interpretable prediction systems capable of handling complex, multimodal healthcare data. This study proposes a novel hybrid framework Causal Graph Attention Network with Large Language Model and Linear Non-Gaussian Acyclic Model (Causal-GAT-LLM-LiNGAM) for robust and explainable heart disease prediction. The approach begins by applying the LiNGAM algorithm to structured clinical data to uncover directed causal relationships among variables such as blood pressure, cholesterol, and smoking history. These causal graphs form the backbone of a Graph Attention Network (GAT), which learns to emphasize the most influential features through attention mechanisms. Simultaneously, a domain-specific Large Language Model (LLM), such as BioBERT or GPT-based variants, is employed to extract semantic embeddings and causal cues from unstructured clinical text, including physician notes and discharge summaries. These text-derived features are integrated into the graph to enrich node and edge representations. By utilizing attention weights and LLM-generated narratives, the combined model not only improves predictive accuracy but also offers human-interpretable reasons for risk assessment. Evaluated on benchmark datasets such as MIMIC-III and UCI Heart Disease, the proposed framework outperforms traditional GNNs and black-box models in terms of accuracy, AUC, and clinical interpretability. This research demonstrates the synergistic potential of causal inference, graph learning, and large language models in developing next-generation diagnostic tools for personalized cardiovascular healthcare.
