article · Nature Journal of Emerging Sciences Technologies and Innovations
Cardiovascular disease risk was evaluated in a cohort of 23,543 hypertensive individuals selected from a broader dataset of 70,000 cardiovascular patients. Three machine-learning approaches were tested: logistic regression, random forest, and gradient boosting. On a dedicated test set, the tuned gradient boosting model delivered the highest accuracy at 78.59 per cent and an area under the receiver operating characteristic curve of 0.6681. In cross-validation, logistic regression achieved a slightly higher mean score of 0.6628. Model explainability analysis revealed that systolic blood pressure, age, and height served as the most influential indicators of risk, with height proving more predictive than body mass index. Lifestyle variables such as smoking, alcohol consumption, and physical exercise offered minimal predictive contribution. Although discriminative performance remained moderate, routinely recorded clinical measurements demonstrated clear utility for patient risk stratification.
Pinpointing which hypertensive patients face the highest risk of cardiovascular disease is difficult in everyday clinical practice. Demonstrating that standard measurements like blood pressure, age, and body size can stratify risk helps clinicians understand the primary drivers of cardiovascular complications. It also clarifies that complex lifestyle factors may add limited predictive value when using routine clinical datasets.
The work points towards clinical decision-support software designed to help healthcare practitioners stratify cardiovascular risk during routine appointments. Because it relies entirely on common clinical indicators, integration into electronic health record systems is technically straightforward. However, with moderate discriminative performance and testing limited to retrospective data analysis, the technology remains at an early, pre-clinical research stage requiring prospective clinical validation before real-world deployment.
AI-generated from the published abstract. Always read the original work before citing.
Hypertension is one of the most important modifiable risk factors for Cardiovascular Disease (CVD), yet identifying which hypertensive patients are at higher risk remains challenging in clinical practice. This study developed and evaluated three machine-learning models: logistic regression, random forest, and Gradient Boosting for CVD risk prediction in a cohort of 23,543 hypertensive patients drawn from a 70,000 patient cardiovascular dataset. After preprocessing, feature engineering, SMOTE-based class balancing, and hyperparameter tuning via randomized search, model performance was assessed on a held-out test set and validated using 5-fold stratified cross-validation with SMOTE correctly nested inside each fold to avoid data leakage. On the test set, tuned Gradient Boosting model achieved the highest accuracy (78.59%) and AUC-ROC (0.6681), outperforming Logistic Regression (0.6633) and Random Forest (0.6508). cross-validation provided a slightly different perspective: Logistic Regression’s mean AUC-ROC (0.6628) edged out Gradient Boosting (0,6609) and Random Forest (0.6383), SHAP analysis on the Gradient Boosting model identified systolic blood pressure, age, and height as the strongest predictors, with height rivaling systolic blood pressure and surpassing BMI a notable difference from Random Forest’s feature importance ranking. Lifestyle factors (smoking, alcohol, physical activity) contributed minimally. These findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive patients, despite their moderate discriminative performance.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.65752/p69ybe82
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.