Background
Cardiovascular disease (CVD) remains a leading cause of global mortality, with current predictive models limited by narrow feature representation, inconsistent benchmarking, and low interpretability, which constrain clinical adoption.
Methods
We developed a comprehensive machine learning framework using the Cardiovascular Diseases Risk Prediction Dataset (n = 308,854, 19 features). Four clinically engineered composite features BMI Risk Category, Lifestyle Risk Score, Comorbidity Burden Index, and Age-BMI Interaction were incorporated. Class imbalance was addressed with SMOTE. Seven classifiers, including Logistic Regression, KNN, Random Forest, XGBoost, Gradient Boosting, LightGBM, and a Stacking Ensemble, were evaluated using 5-fold stratified cross-validation and GridSearchCV optimization. SHAP analysis was employed to interpret model predictions and ensure clinical transparency. External validation was performed on an independent dataset (n = 1529).
Results
The Stacking Ensemble achieved the highest predictive performance, demonstrating superior accuracy, ROC-AUC, and F1-score. Ablation analysis confirmed the incremental value of each engineered feature, with the Comorbidity Burden Index contributing the most. SHAP analysis identified Age Category, General Health, Diabetes, and Smoking History as the key predictors. External validation confirmed robust generalizability across independent datasets. Statistical testing showed significant performance differences among classifiers.
Conclusion
Integrating clinically engineered features, ensemble meta-learning, and explainable AI provides a robust, interpretable, and high-performing framework for CVD risk prediction. The model demonstrates potential for clinical deployment in risk stratification and decision support, contributing to preventive strategies and informed public health interventions.