article · Engineering Applications of Computational Fluid Mechanics
Coagulant dosage modelling has been widely investigated in the literature; however, existing methods often operate as black boxes, failing to provide clear practical explanations for decision-makers. To address this limitation, this study introduces an interpretable framework using SHAP and LIME to highlight the underlying practical problems. Furthermore, this is the first study of its kind in Algeria, bridging a critical regional gap in the literature. Yet, we investigate the application of boosting machine learning models to predict coagulant dose at the Taksebt drinking water treatment plant. Herein we compare between: (i) multiple linear regression (MLR), (ii) random forest regression (RFR), (iii) adaptive boosting (AdaBoost), (iv) categorical boosting (CatBoost), (v) gradient boosted regression trees (GBRT), (vi) Histogram Gradient Boosting (HistGBRT), (vii) Light Gradient Boosting Machine (LightGBM), (viii) Natural Gradient Boosting (NGBoost), and (ix) eXtreme Gradient Boosting (XGBoost). All models were trained using daily measured water quality variables, while aluminium sulphate (Al2(SO4)3·18H2O) was the coagulant dose. The predictive performance was evaluated using root-mean-square error (RMSE), mean absolute error (MAE), Nash-Sutcliffe efficiency (NSE), and the correlation coefficient (R). Global and local model interpretability were conducted using SHAP and LIME, thereby enabling feature importance ranking. To determine whether the observed differences among models were statistically meaningful, two complementary statistical tests were applied: the Diebold-Mariano test and the Kruskal-Walli’s test. The results indicate that CatBoost delivered the most accurate predictions, achieving R, NSE, RMSE, and MAE values of 0.922, 0.849, 2.498, and 1.702 mg/L, respectively, and significantly outperforming the MLR model, which yielded an R of 0.713, a NSE value of 0.507, RMSE of 4.510 mg/L, and MAE of 3.576 mg/L, respectively. Based on SHAP analysis, UV254 was identified as the most influential feature, contributing 24% to the model predictions, whereas COU showed a negligible effect, accounting for only 5%.HighlightsBoosting models for coagulant dosage prediction.CatBoost provided the best numerical performances.UV254 was the most influential raw water quality variable.SHAP and LIME for model interpretability.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1080/19942060.2026.2706865
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.