Predicting Sleep Quality Using Hybrid Stacking Ensemble Of Random Forest And Catboost With Shap Interpretability

DOI: https://doi.org/10.33650/jeecom.v8i2.17445
Authors

(1) * Baitullah - Hakim   (Universitas Islam Negeri Maulana Malik Ibrahim, Malang)
(2)  Muhammad Faisal   (Universitas Islam Negeri Maulana Malik Ibrahim, Malang)  
        Indonesia
(3)  M. Imamudin   (Universitas Islam Negeri Maulana Malik Ibrahim, Malang)  
        Indonesia
(*) Corresponding Author

Abstract


Sleep quality is an important indicator of physical and mental well-being, yet its prediction from lifestyle variables remains challenging because behavioral and physiological factors interact nonlinearly. This study develops a hybrid stacking ensemble that combines Random Forest and CatBoost as base learners and Logistic Regression as the meta-classifier for three-class sleep-quality prediction (Poor, Moderate, Good). The Sleep Health and Lifestyle Dataset contains 374 records, of which only 109 are unique, and is preprocessed through blood-pressure decomposition, derived-feature construction, categorical encoding, and numerical standardization. Models are evaluated with Accuracy, Precision, Recall, F1-Score, and multiclass ROC-AUC under Stratified 10-Fold Cross Validation, both in its standard form and in a duplicate-aware form in which identical records are kept in the same fold. Under the duplicate-aware protocol, Random Forest achieved 97.86% Accuracy, 91.06% macro F1-Score, and 0.9970 ROC-AUC; CatBoost achieved 97.33%, 90.70%, and 0.9930; and Hybrid Stacking achieved 97.06%, 90.51%, and 0.9784. Recall for the minority Poor class (n = 12) was 66.7% for all three models, and stacking did not improve on its base learners. SHAP analysis identifies Stress Level and Sleep Duration as the two dominant features. The results are specific to this small, largely duplicated dataset and do not support clinical use.



Keywords

sleep quality prediction, Random Forest, CatBoost, stacking ensemble, Optuna, SHAP, machine learning



Full Text: PDF



References


K. Ramar et al., “Sleep is essential to health: An American Academy of Sleep Medicine position statement,” Journal of Clinical Sleep Medicine, vol. 17, no. 10, pp. 2115–2119, 2021, doi: 10.5664/jcsm.9476.

M. Fabbri, A. Beracci, M. Martoni, D. Meneo, L. Tonetti, and V. Natale, “Measuring subjective sleep quality: A review,” International Journal of Environmental Research and Public Health, vol. 18, no. 3, 1082, 2021, doi: 10.3390/ijerph18031082.

V. R. K. Sathish, W. L. Woo, and E. S. L. Ho, “Predicting sleeping quality using convolutional neural networks,” arXiv:2204.13584, 2022, doi: 10.48550/arXiv.2204.13584.

S. Ha et al., “Predicting the risk of sleep disorders using a machine learning–based simple questionnaire: Development and validation study,” Journal of Medical Internet Research, vol. 25, e46520, 2023, doi: 10.2196/46520.

H. Han and J. Oh, “Application of various machine learning techniques to predict obstructive sleep apnea syndrome severity,” Scientific Reports, vol. 13, 6379, 2023, doi: 10.1038/s41598-023-33170-7.

A. Bandyopadhyay and C. Goldstein, “Clinical applications of artificial intelligence in sleep medicine: A sleep clinician’s perspective,” Sleep and Breathing, vol. 27, no. 1, pp. 39–55, 2023, doi: 10.1007/s11325-022-02592-4.

R. Alazaidah, G. Samara, M. Aljaidi, M. Haj Qasem, A. Alsarhan, and M. Alshammari, “Potential of machine learning for predicting sleep disorders: A comprehensive analysis of regression and classification models,” Diagnostics, vol. 14, no. 1, p. 27, 2024, doi: 10.3390/diagnostics14010027.

A. B. Mawardi, R. S. Pradini, and M. S. Haris, “Komparasi algoritma boosting untuk prediksi gangguan tidur,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 3, 2025, doi: 10.23960/jitet.v13i3.7281.

A. Zaman, S. Kumar, S. Shatabda, I. Dehzangi, and A. Sharma, “SleepBoost: A multi-level tree-based ensemble model for automatic sleep stage classification,” Medical & Biological Engineering & Computing, vol. 62, no. 9, pp. 2769–2783, 2024, doi: 10.1007/s11517-024-03096-x.

C. Sakal, T. Chen, W. Xu, W. Zhang, Y. Yang, and X. Li, “Towards proactively improving sleep: Machine learning and wearable device data forecast sleep efficiency 4–8 hours before sleep onset,” SLEEP, vol. 48, no. 8, zsaf113, 2025, doi: 10.1093/sleep/zsaf113.

A. Kaur and N. Neeru, “Automated detection of sleep disorders using ensemble learning on clinical and behavioral features,” Sleep Medicine Research, vol. 17, no. 1, pp. 6–20, 2026, doi: 10.17241/smr.2025.03083.

L. Grinsztajn, E. Oyallon, and G. Varoquaux, “Why do tree-based models still outperform deep learning on typical tabular data?,” Advances in Neural Information Processing Systems, vol. 35, 2022, doi: 10.48550/arXiv.2207.08815.

N. Safaei et al., “E-CatBoost: An efficient machine learning framework for predicting ICU mortality using the eICU Collaborative Research Database,” PLoS ONE, vol. 17, no. 5, e0262895, 2022, doi: 10.1371/journal.pone.0262895.

L. Grinsztajn, E. Oyallon, and G. Varoquaux, “Why do tree-based models still outperform deep learning on typical tabular data?,” in Advances in Neural Information Processing Systems, vol. 35, 2022, doi: 10.48550/arXiv.2207.08815.

I. D. Mienye and Y. Sun, “A survey of ensemble learning: Concepts, algorithms, applications, and prospects,” IEEE Access, vol. 10, pp. 99129–99149, 2022, doi: 10.1109/ACCESS.2022.3207287.

H. Chen, S. M. Lundberg, and S.-I. Lee, “Explaining a series of models by propagating Shapley values,” Nature Communications, vol. 13, 4512, 2022, doi: 10.1038/s41467-022-31384-3.

R. Dwivedi et al., “Explainable AI (XAI): Core Ideas, Techniques, and Solutions,” ACM Computing Surveys, vol. 55, no. 9, Article 194, 2023, doi: 10.1145/3561048.

L. Tharmalingam, “Sleep Health and Lifestyle Dataset,” Kaggle. [Online]. Available: https://www.kaggle.com/datasets/uom190346a/sleep-health-and-lifestyle-dataset

B. Bischl et al., “Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges,” WIREs Data Mining and Knowledge Discovery, vol. 13, no. 2, e1484, 2023, doi: 10.1002/widm.1484.

R. D. King, O. I. Orhobor, and C. C. Taylor, “Cross-validation is safe to use,” Nature Machine Intelligence, vol. 3, p. 276, 2021, doi: 10.1038/s42256-021-00332-z.

S. A. Hicks et al., “On evaluation metrics for medical applications of artificial intelligence,” Scientific Reports, vol. 12, 5979, 2022, doi: 10.1038/s41598-022-09954-8.

J. Opitz, “A closer look at classification evaluation metrics and a critical reflection of common evaluation practice,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 820–836, 2024, doi: 10.1162/tacl_a_00675.

Z. Yang et al., “Learning with multiclass AUC: Theory and algorithms,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, pp. 7747–7763, 2022, doi: 10.1109/TPAMI.2021.3101125.

M. Gardani et al., “A systematic review and meta-analysis of poor sleep, insomnia symptoms and stress in undergraduate students,” Sleep Medicine Reviews, vol. 61, 101565, 2022, doi: 10.1016/j.smrv.2021.101565.

S. Atoui et al., “Daily associations between sleep and physical activity: A systematic review and meta-analysis,” Sleep Medicine Reviews, vol. 57, 101426, 2021, doi: 10.1016/j.smrv.2021.101426.

J. Wainer and G. Cawley, “Nested cross-validation when selecting classifiers is overzealous for most practical applications,” Expert Systems with Applications, vol. 182, 115222, 2021, doi: 10.1016/j.eswa.2021.115222.


Dimensions, PlumX, and Google Scholar Metrics

10.33650/jeecom.v8i2.17445


Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Baitullah - Hakim

 
This work is licensed under a Creative Commons Attribution License (CC BY-SA 4.0)

Journal of Electrical Engineering and Computer (JEECOM)
Published by LP3M Nurul Jadid University, Indonesia, Probolinggo, East Java, Indonesia.