Explainable Machine Learning for Predicting Student Dropout and Academic Success Using XGBoost and SHAP

DOI: https://doi.org/10.33650/jeecom.v8i1.16978
Authors

(1) * Sidik Praptomo   (Universitas Muhammadiyah Muara Bungo)  
        Indonesia
(2)  Ahmad Risman   (Universitas Muhammadiyah Muara Bungo)  
        Indonesia
(3)  Riko Muhammad Suri   (Universitas Muhammadiyah Muara Bungo)  
        Indonesia
(*) Corresponding Author

Abstract


Student dropout is a persistent challenge in higher education, and predictive models can support early identification of students who may require academic or financial intervention. This study develops an explainable multiclass machine learning approach to predict three academic outcomes—Dropout, Enrolled, and Graduate—using the public Predict Students' Dropout and Academic Success dataset containing 4,424 student records. Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost) were compared using a stratified 80:20 hold-out design. XGBoost hyperparameters were optimized through randomized search with five-fold stratified cross-validation, and SHapley Additive exPlanations (SHAP) were used to interpret global and class-specific predictions. Random Forest achieved the highest overall accuracy of 77.18%, whereas the optimized XGBoost model produced the highest macro recall of 69.58% and macro F1-score of 70.08%. XGBoost improved recall for the minority Enrolled class to 46.54%, compared with 38.36% for Random Forest and 33.33% for Logistic Regression. SHAP analysis identified the number of curricular units approved in the second and first semesters, tuition-fee status, course, second-semester grade, and age at enrollment among the most influential predictors. Low academic progression and unpaid tuition status contributed strongly toward Dropout predictions, while stronger academic progression shifted predictions toward Graduate. These findings show that explainability complements predictive performance by revealing actionable patterns behind multiclass student-outcome predictions.

Student dropout is a persistent challenge in higher education, and predictive models can support early identification of students who may require academic or financial intervention. This study develops an explainable multiclass machine learning approach to predict three academic outcomes—Dropout, Enrolled, and Graduate—using the public Predict Students' Dropout and Academic Success dataset containing 4,424 student records. Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost) were compared using a stratified 80:20 hold-out design. XGBoost hyperparameters were optimized through randomized search with five-fold stratified cross-validation, and SHapley Additive exPlanations (SHAP) were used to interpret global and class-specific predictions. Random Forest achieved the highest overall accuracy of 77.18%, whereas the optimized XGBoost model produced the highest macro recall of 69.58% and macro F1-score of 70.08%. XGBoost improved recall for the minority Enrolled class to 46.54%, compared with 38.36% for Random Forest and 33.33% for Logistic Regression. SHAP analysis identified the number of curricular units approved in the second and first semesters, tuition-fee status, course, second-semester grade, and age at enrollment among the most influential predictors. Low academic progression and unpaid tuition status contributed strongly toward Dropout predictions, while stronger academic progression shifted predictions toward Graduate. These findings show that explainability complements predictive performance by revealing actionable patterns behind multiclass student-outcome predictions.



Keywords

Educational Data Mining; Student Dropout; XGBoost; Explainable Artificial Intelligence; SHAP



Full Text: PDF



References


V. Realinho, J. Machado, L. Baptista, and M. V. Martins, “Predicting Student Dropout and Academic Success,” Data, vol. 7, no. 11, Art. no. 146, 2022, doi: 10.3390/data7110146.

M. V. Martins, L. Baptista, J. Machado, and V. Realinho, “Multi-Class Phased Prediction of Academic Performance and Dropout in Higher Education,” Applied Sciences, vol. 13, no. 8, Art. no. 4702, 2023, doi: 10.3390/app13084702.

M. Cannistrà, C. Masci, F. Ieva, T. Agasisti, and A. M. Paganoni, “Early-predicting dropout of university students: an application of innovative multilevel machine learning and statistical techniques,” Studies in Higher Education, vol. 47, no. 9, pp. 1935–1956, 2022, doi: 10.1080/03075079.2021.2018415.

J. Niyogisubizo, L. Liao, E. Nziyumva, E. Murwanashyaka, and P. C. Nshimyumukiza, “Predicting student’s dropout in university classes using two-layer ensemble machine learning approach: A novel stacked generalization,” Computers and Education: Artificial Intelligence, vol. 3, Art. no. 100066, 2022, doi: 10.1016/j.caeai.2022.100066.

Z. Song, S.-H. Sung, D.-M. Park, and B.-K. Park, “All-Year Dropout Prediction Modeling and Analysis for University Students,” Applied Sciences, vol. 13, no. 2, Art. no. 1143, 2023, doi: 10.3390/app13021143.

M. Vaarma and H. Li, “Predicting student dropouts with machine learning: An empirical study in Finnish higher education,” Technology in Society, vol. 76, Art. no. 102474, 2024, doi: 10.1016/j.techsoc.2024.102474.

A. Villar and C. R. V. de Andrade, “Supervised machine learning algorithms for predicting student dropout and academic success: a comparative study,” Discover Artificial Intelligence, vol. 4, Art. no. 2, 2024, doi: 10.1007/s44163-023-00079-z.

M. Nagy and R. Molontay, “Interpretable Dropout Prediction: Towards XAI-Based Personalized Intervention,” International Journal of Artificial Intelligence in Education, vol. 34, no. 2, pp. 274–300, 2024, doi: 10.1007/s40593-023-00331-8.

A. Zanellati, S. P. Zingaro, and M. Gabbrielli, “Balancing Performance and Explainability in Academic Dropout Prediction,” IEEE Transactions on Learning Technologies, vol. 17, pp. 2086–2099, 2024, doi: 10.1109/TLT.2024.3425959.

B. Carballo-Mendívil, A. Arellano-González, N. J. Ríos-Vázquez, and M. del P. Lizardi-Duarte, “Predicting Student Dropout from Day One: XGBoost-Based Early Warning System Using Pre-Enrollment Data,” Applied Sciences, vol. 15, no. 16, Art. no. 9202, 2025, doi: 10.3390/app15169202.

H. Khosravi et al., “Explainable Artificial Intelligence in education,” Computers and Education: Artificial Intelligence, vol. 3, Art. no. 100074, 2022, doi: 10.1016/j.caeai.2022.100074.

T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 785–794, doi: 10.1145/2939672.2939785.

S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 4765–4774.

N. Mduma, “Data Balancing Techniques for Predicting Student Dropout Using Machine Learning,” Data, vol. 8, no. 3, Art. no. 49, 2023, doi: 10.3390/data8030049.


Dimensions, PlumX, and Google Scholar Metrics

10.33650/jeecom.v8i1.16978


Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Sidik Praptomo, Ahmad Risman, Riko Muhammad Suri

 
This work is licensed under a Creative Commons Attribution License (CC BY-SA 4.0)

Journal of Electrical Engineering and Computer (JEECOM)
Published by LP3M Nurul Jadid University, Indonesia, Probolinggo, East Java, Indonesia.