article
The application of machine learning (ML) to medical datasets offers significant potential for improving disease prediction and patient outcomes. However, challenges such as feature redundancy, overfitting, and suboptimal model performance limit the practical effectiveness of ML algorithms. This study focuses on optimizing ML techniques for cardiovascular disease prediction using the Kaggle Cardiovascular Disease dataset. We systematically apply feature selection methods, including correlation analysis and regularization techniques (L1/L2), to identify the most relevant attributes and address multicollinearity. Advanced ensemble models such as Random Forest, XGBoost, and LightGBM are employed to mitigate overfitting and enhance predictive performance. Through hyperparameter tuning and stratified k-fold cross-validation, we ensure model robustness and generalizability. The results demonstrate that ensemble methods, particularly gradient boosting algorithms, outperform traditional models, achieving superior predictive accuracy and stability. This study highlights the importance of algorithm optimization in ML applications for healthcare, offering a replicable framework for medical datasets and paving the way for more effective diagnostic tools in cardiovascular health.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.3390/cmsf2025010013
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.