From Black Box to Grid-Ready: Explainable CRISP-DM for High-Demand EV Charging Prediction

DOI: https://doi.org/10.33650/jeecom.v8i2.17161
Authors

(1) * Dwi Putro Sarwo Setyohadi   (Politeknik Negeri Jember)
(2)  Hendra Yufit Riskiawan   (Politeknik Negeri Jember)
(3)  Akas Bagus Setiawan   (Politeknik Negeri Jember)
(*) Corresponding Author

Abstract


The rapid growth of electric vehicle (EV) adoption creates uncertainty for power grid operators in anticipating charging load, particularly sessions that draw a disproportionately large amount of energy. This study applies the CRISP-DM (Cross-Industry Standard Process for Data Mining) framework to a large-scale synthetic EV charging dataset (1,965,239 sessions; four raw attributes: connection time, charging duration, energy delivered, and a sequential day index) to build an early, context-only classifier for high-demand sessions, defined as sessions whose delivered energy falls in the top quartile (≥75th percentile). To avoid label leakage, only information available at the start of a session connection time, planned duration, a peak-hour flag, and a weekend flag was used as predictors. The resulting target is naturally imbalanced (75:25); its effect was investigated through an ablation study comparing five classification algorithms (Logistic Regression, Decision Tree, Random Forest, XGBoost, and K-Nearest Neighbors) with and without SMOTE oversampling, evaluated on accuracy, precision, recall, F1-score, and ROC-AUC. The best configuration, XGBoost combined with SMOTE, achieved F1-score = 0.600 and ROC-AUC = 0.818, improving recall from 0.442 to 0.807 at a moderate precision cost. SHAP-based attribution identified charging duration as the dominant predictor in the trained model, followed by connection time, consistent with the raw feature correlations (r = 0.51 and r = 0.30, respectively). The full pipeline data preparation, modeling, ablation, and explainability were deployed as an interactive, open web application built with Streamlit, allowing non-technical stakeholders to reproduce the analysis and inspect results in real time.


Keywords

Class Imbalance; CRISP-DM; Electric Vehicle Charging; Explainable AI; SMOTE; Streamlit; XGBoost



Full Text: PDF



References


S. Ghanbari Motlagh and L. Li, “A review on electric vehicle charging station planning: Infrastructure placement, sizing, upgrades, and uncertainties,” J. Energy Storage, vol. 141, p. 119325, 2026, doi: 10.1016/j.est.2025.119325.

S. Mohanty et al., “Demand side management of electric vehicles in smart grids: A survey on strategies, challenges, modeling, and optimization,” Energy Reports, vol. 8, pp. 12466–12490, 2022, doi: 10.1016/j.egyr.2022.09.023.

S. Fatima, V. Püvi, M. Lehtonen, and M. Pourakbari-Kasmaei, “A review of electric vehicle hosting capacity quantification and improvement techniques for distribution networks,” IET Gener. Transm. Distrib., vol. 18, no. 6, pp. 1095–1113, 2024, doi: 10.1049/gtd2.13010.

B. Palaniyappan and T. Vinopraba, “Dynamic pricing for load shifting: Reducing electric vehicle charging impacts on the grid through machine learning-based demand response,” Sustain. Cities Soc., vol. 103, p. 105256, 2024, doi: 10.1016/j.scs.2024.105256.

G. Vishnu, D. Kaliyaperumal, P. B. Pati, A. Karthick, N. Subbanna, and A. Ghosh, “Short-term forecasting of electric vehicle load using time series, machine learning, and deep learning techniques,” World Electr. Veh. J., vol. 14, no. 9, p. 266, 2023, doi: 10.3390/wevj14090266.

I. Ullah, K. Liu, T. Yamamoto, M. Zahid, and A. Jamal, “Modeling of machine learning with SHAP approach for electric vehicle charging station choice behavior prediction,” Travel Behav. Soc., vol. 31, pp. 78–92, 2023, doi: 10.1016/j.tbs.2022.11.006.

T. Zhang, Q. Peng, and S. Zeng, “Predicting EV charging demand in renewable-energy-powered grids using explainable machine learning,” Sustainability, vol. 17, no. 9, p. 4158, 2025, doi: 10.3390/su17094158.

R. Alsaigh, R. Mehmood, and I. Katib, “AI explainability and governance in smart energy systems: A review,” Front. Energy Res., vol. 11, p. 1071291, 2023, doi: 10.3389/fenrg.2023.1071291.

W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, “A survey on imbalanced learning: Latest research, applications and future directions,” Artif. Intell. Rev., vol. 57, p. 137, 2024, doi: 10.1007/s10462-024-10759-6.

A. Arafa, N. El-Fishawy, M. Badawy, and M. Radad, “RN-SMOTE: Reduced noise SMOTE based on DBSCAN for enhancing imbalanced data classification,” J. King Saud Univ. - Comput. Inf. Sci., vol. 34, no. 8, pp. 5059–5074, 2022, doi: 10.1016/j.jksuci.2022.06.005.

A. Saranya and R. Subhashini, “A systematic review of explainable artificial intelligence models and applications: Recent developments and future trends,” Decis. Anal. J., vol. 7, p. 100230, 2023, doi: 10.1016/j.dajour.2023.100230.

A. V Ponce-Bobadilla, V. Schmitt, C. S. Maier, S. Mensing, and S. Stodtmann, “Practical guide to SHAP analysis: Explaining supervised machine learning model predictions in drug development,” Clin. Transl. Sci., vol. 17, no. 11, p. e70056, 2024, doi: 10.1111/cts.70056.

V. Plotnikova, M. Dumas, and F. P. Milani, “Applying the CRISP-DM data mining process in the financial services industry: Elicitation of adaptation requirements,” Data Knowl. Eng., vol. 139, p. 102013, 2022, doi: 10.1016/j.datak.2022.102013.

G. Douzas, “imbalanced-learn-extra: A Python package for novel oversampling algorithms,” J. Open Res. Softw., 2026, doi: 10.5334/jors.459.

D. Kreuzberger, N. Kühl, and S. Hirschl, “Machine learning operations (MLOps): Overview, definition, and architecture,” IEEE Access, vol. 11, pp. 31866–31879, 2023, doi: 10.1109/access.2023.3262138.

M. Zarour, H. Alzabut, and K. T. Al-Sarayreh, “MLOps best practices, challenges and maturity models: A systematic literature review,” Inf. Softw. Technol., vol. 183, p. 107733, 2025, doi: 10.1016/j.infsof.2025.107733.

L. Pilgram, H. Ko, A. Tung, and K. El Emam, “Protecting patient privacy in tabular synthetic health data: A regulatory perspective,” npj Digit. Med., vol. 8, p. 732, 2025, doi: 10.1038/s41746-025-02112-0.

Y. Zhang et al., “A high-resolution electric vehicle charging transaction dataset with multidimensional features in China,” Sci. Data, vol. 12, p. 643, 2025, doi: 10.1038/s41597-025-04982-1.

S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, p. 100804, 2023, doi: 10.1016/j.patter.2023.100804.

S. Somvanshi, S. Das, S. Javed, G. Antariksa, and A. Hossain, “A survey on tabular data: From tree-based methods to tabular deep learning,” ACM Comput. Surv., 2026, doi: 10.1145/3807777.

F. E. Arévalo-Cordovilla and M. Peña, “Comparative analysis of machine learning models for predicting student success in online programming courses: A study based on LMS data and external factors,” Mathematics, vol. 12, no. 20, p. 3272, 2024, doi: 10.3390/math12203272.

V. G. Costa and C. E. Pedreira, “Recent advances in decision trees: An updated survey,” Artif. Intell. Rev., vol. 56, pp. 4765–4800, 2023, doi: 10.1007/s10462-022-10275-5.

M. Aria, C. Cuccurullo, and A. Gnasso, “A comparison among interpretative proposals for random forests,” Mach. Learn. with Appl., vol. 6, p. 100094, 2021, doi: 10.1016/j.mlwa.2021.100094.

G. Velarde et al., “Tree boosting methods for balanced and imbalanced classification and their robustness over time in risk assessment,” Mach. Learn. with Appl., 2025, doi: 10.1016/j.iswa.2024.200354.

A. Ali, M. Hamraz, N. Gul, D. M. Khan, S. Aldahmani, and Z. Khan, “A k nearest neighbour ensemble via extended neighbourhood rule and feature subsets,” Pattern Recognit., vol. 142, p. 109641, 2023, doi: 10.1016/j.patcog.2023.109641.

K. Hamad, L. Obaid, S. Haridy, W. Zeiada, and G. G. Al-Khateeb, “Factorial design-machine learning approach for predicting incident durations,” Comput. Civ. Infrastruct. Eng., vol. 38, no. 5, pp. 660–680, 2023, doi: 10.1111/mice.12883.

A. F. A. Sayed, M. A. Arafa, N. A. El-Nimr, K. A. S. Banawan, and M. S. Abdou, “Comparison of machine learning classification and regression models for prediction of academic performance among postgraduate public health students,” Sci. Rep., vol. 15, p. 44056, 2025, doi: 10.1038/s41598-025-31023-z.

T. W. Campbell, H. Roder, R. W. Georgantas III, and J. Roder, “Exact Shapley values for local and model-true explanations of decision tree ensembles,” Mach. Learn. with Appl., vol. 9, p. 100345, 2022, doi: 10.1016/j.mlwa.2022.100345.

A. B. Setiawan, H. Y. Riskiawan, H. A. Putranto, T. Rizaldi, and R. A. Atmoko, “An end-to-end machine learning pipeline for online purchase intention prediction using Random Forest and MLOps practices,” Angkasa J. Ilm. Bid. Teknol., vol. 18, no. 1, pp. 10–20, 2026, doi: 10.28989/angkasa.v18i1.3841.

S. P. Revathy, M. Sindhuja, and R. Jayashree, “Streamlit-based web application for Parkinson’s detection using machine learning,” J. Artif. Intell. Capsul. Networks, vol. 6, no. 4, pp. 415–428, 2024, doi: 10.36548/jaicn.2024.4.006.

A. R. Singh, R. S. Kumar, M. Bajaj, C. B. Khadse, and I. Zaitsev, “Machine learning-based energy management and power forecasting in grid-connected microgrids with multiple distributed energy sources,” Sci. Rep., vol. 14, 2024, doi: 10.1038/s41598-024-70336-3.

A. Lavanya, S. Ali, and S. D. Vidya Sagar, “Assessing the performance of Python data visualization libraries: A review,” Int. J. Comput. Eng. Res. Trends, vol. 10, no. 1, pp. 29–39, 2023, doi: 10.22362/ijcert/2023/v10/i01/v10i0104.

F. Nahrstedt, M. Karmouche, K. Bargieł, P. Banijamali, A. N. Pradeep Kumar, and I. Malavolta, “An empirical study on the energy usage and performance of Pandas and Polars data analysis Python libraries,” 2024, Salerno, Italy. doi: 10.1145/3661167.3661203.

A. B. Setiawan, H. Y. Riskiawan, T. Rizaldi, H. A. Putranto, R. A. Atmoko, A. B. F. Mansur, M. Alharthi, and A. H. Basori, “From chaos to Gleichgewicht: An application-specific BiLSTM framework with SMOTE, SHAP, and TF-IDF-based safety support for proxy risk screening on noisy social media,” Computation, vol. 14, no. 8, art. 167, 2026, https://doi.org/10.3390/computation14080167.

D. P. S. Setyohadi, H. Y. Riskiawan, A. S. Arifianto, I. G. Wiryawan, and A. B. Setiawan, “Explainable clinical-operational intelligence for hospital length of stay prediction using integrated multi-source admission data with time-based evaluation,” J. Vocational Inform. Comput. Educ., vol. 4, no. 2, 2026, https://doi.org/10.66053/voice.v4i2.507.


Dimensions, PlumX, and Google Scholar Metrics

10.33650/jeecom.v8i2.17161


Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Dwi Putro Sarwo Setyohadi, Hendra Yufit Riskiawan, Akas Bagus Setiawan

 
This work is licensed under a Creative Commons Attribution License (CC BY-SA 4.0)

Journal of Electrical Engineering and Computer (JEECOM)
Published by LP3M Nurul Jadid University, Indonesia, Probolinggo, East Java, Indonesia.