article · Obesities
Obesity is ranked as one of the biggest health challenges facing humanity today. Globally, the number of obese people has almost tripled since 1975, and this lifestyle disease currently affects hundreds of millions of adults who suffer from major health problems due to it, such as heart disease, type 2 diabetes and some cancers, that weigh heavily on the global health systems, In order to keep high standards for methods, anthropometric variables, i.e., Height and Weight have been intentionally excluded from the features, because labels for obesity classes are based on these measurements; thus, including them would introduce target leakage. All models were individually tuned with Optuna (50 trials, TPE sampler), and the class imbalance was managed by the synthetic minority over-sampling technique (SMOTE), which was done only in training folds. The models were evaluated by stratified five-fold cross-validation, with the macro-averaged F1-score being used as the main metric for evaluation. The best model was the fine-tuned XGBoost, which gave a test macro F1-score value of 0.872 and a macro-AUC of 0.977. The model was higher performing than others such as Random Forest (F1 = 0.869), MLP (F1 = 0.777), and Logistic Regression (F1 = 0.605). This means that behavioral and lifestyle variables may have a very strong and sufficient signal to identify obesity status, even when there are no direct anthropometric measurements available. However, it is worth noting that results here represent only performance on a single public benchmark dataset, so they cannot be taken as proof that the model would do well in real-world clinical settings. With the advent of ML methods for obesity prediction, rigorous, leakage-free evaluation becomes indispensable. Apart from external validation of the clinical models on independent datasets, the use of interpretability tools such as SHapley Additive exPlanations (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME) for understanding decision-making, as well as sex and gender subgroup analyses for evaluating fairness and equity, should also be pursued in the future. This study highlights the importance of rigorous, leakage-free evaluation in machine learning-based obesity research. Future work should focus on external validation using independent clinical cohorts, the integration of interpretability techniques such as SHapley Additive exPlanations (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME), and subgroup analyses by sex and gender to assess model fairness and clinical equity.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.3390/obesities6030027
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.