MARATTO

article · Discover Artificial Intelligence

Transparent machine learning-based classification of cholesterol levels using random forest and ELI5 explainability

2026Open accessGondar University

Abstract

Elevated low-density lipoprotein cholesterol (LDL-C) is a major risk factor for cardiovascular disease, yet conventional clinical thresholds often fail to capture complex interactions among demographic, metabolic, and behavioral factors. Machine-learning (ML) approaches offer improved classification and risk stratification beyond traditional cutoffs, particularly when combined with explainable methods. We analyzed WHO STEPS survey data (2018 onward) from low- and lower-middle-income countries, including 61,780 adults (44,782 normal and 16,998 abnormal cholesterol). Missing values were imputed using K-nearest neighbors, and categorical variables were numerically encoded. Recursive Feature Elimination with cross-validation (RFECV) using Random Forest identified the top 15 predictors. Nine ML classifiers were trained on an 80/20 train–test split and evaluated using accuracy, precision, recall, F1-score, and area under the ROC curve (AUC). Model interpretability was assessed using permutation feature importance implemented through ELI5. Recursive feature elimination feature selection method highlighted metabolic and socio-demographic factors including alcohol consumption, fasting blood sugar, age group, weight, education level, marital status, and residency as the most influential predictors of abnormal cholesterol. Ensemble models consistently outperformed traditional classifiers. Random Forest achieved the best performance (accuracy = 96.66%, F1-score = 96.58%, AUC = 0.993), demonstrating strong discriminatory ability across precision–recall thresholds. Permutation analysis further identified hypertension status, alcohol use, physical activity, waist category, and residency as key determinants. Explainable ensemble-based ML models effectively capture the multifactorial determinants of cholesterol abnormality in population-level data. This transparent approach offers a robust and practical tool for early cardiovascular risk stratification in resource-limited settings.

Research topics

  • Artificial Intelligence in Healthcare
  • Diabetes, Cardiovascular Risks, and Lipoproteins
  • Lipoproteins and Cardiovascular Health

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1007/s44163-026-00974-1

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.