article
The Programme for International Student Assessment (PISA) serves as a global benchmark for evaluating the academic performance of 15-year-old students across mathematics, reading, and science. This study applies interpretable machine learning techniques to analyze 2,087 PISA records from 47 countries spanning 2000 to 2018. We compare the predictive performance of three models XGBoost, CatBoost, and Random Forest on PISA score prediction. Among them, XGBoost achieved the lowest test RMSE (7.60), outperforming CatBoost (12.24) and Random Forest (24.62). To ensure interpretability, we used SHAP (SHapley Additive exPlanations) with the XGBoost model, which supports efficient integration with the TreeExplainer algorithm, enabling transparent attribution of predictions to input features. SHAP analysis reveals that country-level context overwhelmingly shapes educational outcomes, far outweighing subject area, gender, or year. Visualizations further highlight stark cross-national disparities, with nations like Singapore and Hong Kong contributing over 45 points to student performance, while countries such as Brazil and Mexico exhibit substantial deficits. These insights emphasize the persistent and systemic nature of global educational inequality and demonstrate the utility of combining high-performance prediction with interpretable machine learning for educational research and policy-making.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/commnet68224.2025.11288901
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.