Optimization of Machine Learning Algorithms in Breast Cancer Classification: A Performance Based Analysis

DOI: https://doi.org/10.33650/jeecom.v8i1.17160
Authors

(1)  Agus Wantoro   (Universitas Aisyah Pringsewu)  
        Indonesia
(2) * Arie Setya Putra   (Universitas Mitra Indonesia)  
        Indonesia
(3)  Ochi Marshella Febriani   (Institut Informatika dan Bisnis Darmajaya)  
        Indonesia
(*) Corresponding Author

Abstract


Timely identification of breast cancer recurrence is closely associated with patient survival and the effectiveness of treatment. Inaccurate detection can contribute to greater disease severity, higher treatment costs, longer recovery, and reduced quality of care. For Machine Learning (ML)-based decision-support systems, two important challenges are the unequal distribution of medical-data classes and the large number of features, both of which may affect model accuracy and computational efficiency. This study evaluates an approach that combines feature selection with class-imbalance handling to improve breast cancer detection performance. Information Gain (IG), Gain Ratio (GR), Gini Decrease (GD), and Relief-F are used to rank features according to their weights, while the Synthetic Minority Over-Sampling Technique (SMOTE) is applied to improve representation of the underrepresented class. Seven ML classifiers, namely k-Nearest Neighbor (k-NN), Tree, Support Vector Machine (SVM), Naive Bayes, AdaBoost, Random Forest (RF), and Neural Network (NN), are tested and assessed using confusion-matrix-based accuracy, precision, recall, and computational time. The experimental results indicate that incorporating class-imbalance handling improves the predictive performance of the ML algorithms. Among the evaluated combinations, Information Gain with Random Forest (IG+RF) provides the optimal result in this case. These findings highlight the value of integrating class-balancing and feature-selection procedures when developing machine-learning systems for breast cancer detection.


Keywords

Machine Learning; Imbalance Class; Feature Selection; Breast Cancer; Classification;



Full Text: PDF



References


A. Schindele et al., “Interpretable machine learning for thyroid cancer recurrence predicton: Leveraging XGBoost and SHAP analysis,” Eur. J. Radiol., vol. 186, May 2025, doi: 10.1016/j.ejrad.2025.112049.

A. H. Barfejani et al., “Predicting overall survival in anaplastic thyroid cancer using machine learning approaches,” Eur. Arch. Oto-Rhino-Laryngology, vol. 282, no. 3, pp. 1653–1657, 2025, doi: 10.1007/s00405-024-08986-2.

D. W. Chen, B. H. H. Lang, D. S. A. McLeod, K. Newbold, and M. R. Haymart, “Thyroid cancer,” Lancet, vol. 401, no. 10387, pp. 1531–1544, May 2023, doi: 10.1016/S0140-6736(23)00020-X.

A. Kuang, V. L. Kouznetsova, S. Kesari, and I. F. Tsigelny, “Diagnostics of Thyroid Cancer Using Machine Learning and Metabolomics,” Metabolites, vol. 14, no. 1, 2024, doi: 10.3390/metabo14010011.

R. Iacob et al., “Evaluating the Role of Breast Ultrasound in Early Detection of Breast Cancer in Low- and Middle-Income Countries: A Comprehensive Narrative Review,” 2024. doi: 10.3390/bioengineering11030262.

Y.-M. Huang et al., “Correction: Huang et al. Systemic Anticoagulation and Inpatient Outcomes of Pancreatic Cancer: Real-World Evidence from U.S. Nationwide Inpatient Sample. Cancers 2023, 15, 1985,” 2024. doi: 10.3390/cancers16061181.

I. O. Lixandru-Petre et al., “Machine Learning for Thyroid Cancer Detection, Presence of Metastasis, and Recurrence Predictions—A Scoping Review,” Cancers (Basel)., vol. 17, no. 8, pp. 1–27, 2025, doi: 10.3390/cancers17081308.

S. Li, Z. Tang, L. Yang, M. Li, and Z. Shang, “Application of deep reinforcement learning for spike sorting under multi-class imbalance,” Comput. Biol. Med., vol. 164, p. 107253, 2023, doi: https://doi.org/10.1016/j.compbiomed.2023.107253.

X. Song et al., “Evolutionary computation for feature selection in classification: A comprehensive survey of solutions, applications and challenges,” Swarm Evol. Comput., vol. 90, p. 101661, 2024, doi: https://doi.org/10.1016/j.swevo.2024.101661.

W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, “A survey on imbalanced learning: latest research, applications and future directions,” Artif. Intell. Rev., vol. 57, no. 6, p. 137, 2024, doi: 10.1007/s10462-024-10759-6.

L. C. M. Liaw, S. C. Tan, P. Y. Goh, and C. P. Lim, “A histogram SMOTE-based sampling algorithm with incremental learning for imbalanced data classification,” Inf. Sci. (Ny)., vol. 686, p. 121193, 2025, doi: https://doi.org/10.1016/j.ins.2024.121193.

K. E. Setiawan, “Predicting Recurrence in Differentiated Thyroid Cancer: a Comparative Analysis of Various Machine Learning Models Including Ensemble Methods With Chi-Squared Feature Selection,” Commun. Math. Biol. Neurosci., vol. 2024, no. Scenario 1, pp. 1–29, 2024, doi: 10.28919/cmbn/8506.

M. S. H. Shaon, T. Karim, M. S. Shakil, and M. Z. Hasan, “A comparative study of machine learning models with LASSO and SHAP feature selection for breast cancer prediction,” Healthc. Anal., vol. 6, p. 100353, 2024, doi: https://doi.org/10.1016/j.health.2024.100353.

S. Benghazouani, S. Nouh, and A. Zakrani, “Optimizing breast cancer diagnosis: Harnessing the power of nature-inspired metaheuristics for feature selection with soft voting classifiers,” Int. J. Cogn. Comput. Eng., vol. 6, pp. 1–20, 2025, doi: https://doi.org/10.1016/j.ijcce.2024.09.005.

S. Batool and S. Zainab, “A comparative performance assessment of artificial intelligence based classifiers and optimized feature reduction technique for breast cancer diagnosis,” Comput. Biol. Med., vol. 183, p. 109215, 2024, doi: https://doi.org/10.1016/j.compbiomed.2024.109215.

B. Zielosko and A. Dmytrenko, “Selected Methods of Feature Selection - Medical Case Study,” Procedia Comput. Sci., vol. 246, pp. 4451–4460, 2024, doi: https://doi.org/10.1016/j.procs.2024.09.295.

K. Premalatha, D. Prabha Devi, and K. Sivakumar, “Machine learning framework for breast cancer detection with feature selection with L2 ridge regularization: insights from multiple datasets,” J. Transl. Genet. Genomics, vol. 9, no. 1, pp. 11–34, 2025, doi: 10.20517/jtgg.2024.82.

I. Chhillar and A. Singh, “An improved soft voting-based machine learning technique to detect breast cancer utilizing effective feature selection and SMOTE-ENN class balancing,” Discov. Artif. Intell., vol. 5, no. 1, p. 4, 2025, doi: 10.1007/s44163-025-00224-w.

D. S. Hassan, “The effect of feature selection methods on machine learning model performance: a comparative study for breast cancer prediction,” Sci. J. Univ. Zakho, vol. 13, no. 1, pp. 102–113, 2025, doi: 10.25271/sjuoz.2025.13.1.1429.

A. Yaqoob et al., “SGA-Driven feature selection and random forest classification for enhanced breast cancer diagnosis: A comparative study.,” Sci. Rep., vol. 15, no. 1, p. 10944, Mar. 2025, doi: 10.1038/s41598-025-95786-1.

G. Alfian et al., “Predicting Breast Cancer from Risk Factors Using SVM and Extra-Trees-Based Feature Selection Method,” 2022. doi: 10.3390/computers11090136.

W. M. Shaban, “Insight into breast cancer detection: new hybrid feature selection method,” Neural Comput. Appl., vol. 35, no. 9, pp. 6831–6853, 2023, doi: 10.1007/s00521-022-08062-y.

K. Kannadasan, D. R. Edla, and V. Kuppili, “Type 2 diabetes data classification using stacked autoencoders in deep neural networks,” Clin. Epidemiol. Glob. Heal., vol. 7, no. 4, pp. 530–535, 2019, doi: 10.1016/j.cegh.2018.12.004.

C. Sharma and A. Singla, “Advanced PTSVM Based Breast Cancer Classification with Weighted Feature Selection,” SN Comput. Sci., vol. 6, no. 1, p. 50, 2024, doi: 10.1007/s42979-024-03590-x.

G. Husain et al., “SMOTE vs. SMOTEENN: A Study on the Performance of Resampling Algorithms for Addressing Class Imbalance in Regression Models,” Algorithms, vol. 18, no. 1, pp. 1–16, 2025, doi: 10.3390/a18010037.

M. F. Ijaz, G. Alfian, M. Syafrudin, and J. Rhee, “Hybrid Prediction Model for type 2 diabetes and hypertension using DBSCAN-based outlier detection, Synthetic Minority Over Sampling Technique (SMOTE), and random forest,” Appl. Sci., vol. 8, no. 8, 2018, doi: 10.3390/app8081325.

H. Sulistiani, A. Syarif, K. Muludi, and Warsito, “Performance evaluation of feature selections on some ML approaches for diagnosing the narcissistic personality disorder,” Bull. Electr. Eng. Informatics, vol. 13, no. 2, pp. 1383–1391, 2024, doi: 10.11591/eei.v13i2.6717.

J. Wang, S. Zhou, Y. Yi, and J. Kong, “An improved feature selection based on effective range for classification,” Sci. World J., vol. 2014, 2014, doi: 10.1155/2014/972125.

S. Bashir, Z. S. Khan, F. H. Khan, A. Anjum, and K. Bashir, “Improving Heart Disease Prediction Using Feature Selection Approaches,” in 2019 16th International Bhurban Conference on Applied Sciences and Technology (IBCAST), 2019, pp. 619–623. doi: 10.1109/IBCAST.2019.8667106.

J. Gao, Z. Wang, T. Jin, J. Cheng, Z. Lei, and S. Gao, “Information gain ratio-based subfeature grouping empowers particle swarm optimization for feature selection,” Knowledge-Based Syst., vol. 286, p. 111380, 2024, doi: https://doi.org/10.1016/j.knosys.2024.111380.

P. Bhat and K. Dutta, “A multi-tiered feature selection model for android malware detection based on Feature discrimination and Information Gain,” J. King Saud Univ. - Comput. Inf. Sci., vol. 34, no. 10, Part B, pp. 9464–9477, 2022, doi: https://doi.org/10.1016/j.jksuci.2021.11.004.

M. Trabelsi, N. Meddouri, and M. Maddouri, “A New Feature Selection Method for Nominal Classifier based on Formal Concept Analysis,” Procedia Comput. Sci., vol. 112, pp. 186–194, 2017, doi: 10.1016/j.procs.2017.08.227.

E. A. Algehyne, M. L. Jibril, N. A. Algehainy, O. A. Alamri, and A. K. Alzahrani, “Fuzzy Neural Network Expert System with an Improved Gini Index Random Forest-Based Feature Importance Measure Algorithm for Early Diagnosis of Breast Cancer in Saudi Arabia,” 2022. doi: 10.3390/bdcc6010013.

S. Tangirala, “Evaluating the Impact of GINI Index and Information Gain on Classification using Decision Tree Classifier Algorithm *,” Int. J. Adv. Comput. Sci. Appl., vol. 11, no. 2, pp. 612–619, 2020.

Alok Kumar Shukla, Sanjeev Kumar Pippal, Srishti Gupta, B. Ramachandra Reddy, and Diwakar Tripathi, “Knowledge discovery in medical and biological datasets by integration of Relief-F and correlation feature selection techniques,” J. Intell. Fuzzy Syst., vol. 38, no. 5, pp. 6637–6648, May 2020, doi: 10.3233/JIFS-179743.

P. Cunningham and S. J. Delany, “k-Nearest Neighbour Classifiers - A Tutorial,” ACM Comput. Surv., vol. 54, no. 6, Jul. 2021, doi: 10.1145/3459665.

T. Yan, S.-L. Shen, A. Zhou, and X. Chen, “Prediction of geological characteristics from shield operational parameters by integrating grid search and K-fold cross validation into stacking classification algorithm,” J. Rock Mech. Geotech. Eng., vol. 14, no. 4, pp. 1292–1303, 2022, doi: https://doi.org/10.1016/j.jrmge.2022.03.002.

M. Ohsaki, P. Wang, K. Matsuda, S. Katagiri, H. Watanabe, and A. Ralescu, “Confusion-matrix-based kernel logistic regression for imbalanced data classification,” IEEE Trans. Knowl. Data Eng., vol. 29, no. 9, pp. 1806–1819, 2017, doi: 10.1109/TKDE.2017.2682249


Dimensions, PlumX, and Google Scholar Metrics

10.33650/jeecom.v8i1.17160


Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Agus Wantoro

 
This work is licensed under a Creative Commons Attribution License (CC BY-SA 4.0)

Journal of Electrical Engineering and Computer (JEECOM)
Published by LP3M Nurul Jadid University, Indonesia, Probolinggo, East Java, Indonesia.