article
In machine learning-based medical diagnostics, feature selection is an important task. It enhances machine learning models by improving interpretability, diminishing dimensionality, and restricting overfitting. This study investigates the effect of several feature selection methods on heart disease prediction on the Cleveland Heart Disease dataset. Namely: Filter, wrapper, embedded and model-agnostic feature selection. Procedures are assessed together with a number of machine learning classifiers, such as Logistic Regression, Random Forest, Support Vector Machines, K-Nearest Neighbors and XG- Boost. Accuracy, precision, recall, F1-score, and ROC-AUC are used to measure model performance. Findings show that the performance of feature selection methods is not always more optimal compared to using all features. However, these methods enable the use of smaller and clinically meaningful feature subsets with a negligible reduction in the accuracy of classification. Among the classifiers, Random Forest, XGBoost, and Logistic Regression prove strong evidence of performance in various feature selection strategies. The results point out that feature selection is predominantly useful to simplify the model and improve interpretability instead of maximizing the accuracy.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset68627.2026.11538603
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.