article · Physics and Chemistry of the Earth Parts A/B/C
This study presents a machine learning framework for predicting and assessing membrane fouling risk in water treatment plants in Namibia. The approach integrates systematic data preprocessing with advanced predictive modelling techniques to estimate fouling propensity under varying physicochemical and operational conditions. Laboratory water quality data from Grunau, Opuwo, and Eenhana treatment plants were compiled and analysed using Decision Tree, Random Forest, Gradient Boosting, k Nearest Neighbours, and Support Vector Regression models. A continuous fouling severity index was developed, predicted, and analysed using time-series techniques to examine temporal patterns and site-specific dynamics. Model performance was evaluated using MAE, RMSE, R-squared, NSE, KGE, and PBIAS, supported by error distribution analysis and statistical testing. The results show that tree-based models consistently outperform other approaches, with Gradient Boosting achieving the best overall performance, with an R-squared of 0.962, an MAE of 0.793, and strong agreement across all evaluation metrics. Error distribution analysis confirmed that these models provide more stable and consistent predictions, while k-nearest neighbours and Support Vector Regression exhibit higher variability and reduced reliability. The findings indicate that early-stage fouling risk is primarily influenced by Total Dissolved Solids and pH, while breakthrough fouling events are associated with transient increases in iron and manganese concentrations. Time series analysis further revealed site-specific fouling cycles and temporal anomalies across treatment stages. The study demonstrates that tree-based machine learning models provide a reliable and practical tool for predicting fouling risk, supporting improved operational decision-making and proactive management in water treatment systems. • Tree-based models consistently outperformed other approaches, showing strong capability in capturing nonlinear and threshold-driven fouling behaviour. • Gradient Boosting achieved the best overall performance across all evaluation metrics, showing high accuracy, strong agreement, and stable error distribution. • Fouling dynamics are primarily controlled by Total Dissolved Solids and pH at early stages, while transient iron and manganese spikes drive breakthrough fouling events. • Error distribution and statistical tests confirmed significant differences in model performance, with tree-based models providing more reliable and consistent predictions. • Time series analysis revealed site-specific fouling patterns, including reduced scaling trends in Grunau, recurring metal-driven peaks in Opuwo, and consistently low fouling in Eenhana.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.pce.2026.104498
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.