Home / Current Issue / Paper 1711021
Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
DOI: 10.64388/IREV9I4-1711021-6791
Abstract
This study conducts a comparative analysis of three prominent machine learning models?Logistic Regression, Random Forest, and Support Vector Machine (SVM),?for the classification of breast cancer. The data used for this study is collected from the records of incoming patients of breast cancer in Ekiti State University Teaching Hospital, Ekiti State, Nigeria. The analysis emphasizes the significance of certain features, particularly the worst measurements of texture, radius, and area, in distinguishing between benign and malignant tumors. Key features such as `area worst`, `radius worst`, and `texture worst` demonstrate high phi values, indicating their strong association with the class labels. This suggests that extreme values of these measurements are crucial in identifying malignancies. Among the evaluated models, the Random Forest model exhibits the highest accuracy, as validated by cross-validation techniques. The selection of `mtry = 2` as the optimal parameter underscores the importance of choosing the appropriate number of features at each split to maximize model performance. The model's reliability is further confirmed by confusion matrices, which show high sensitivity and specificity, critical for minimizing false negatives and positives in medical diagnoses. This study highlights the importance of feature importance analysis in medical data classification, revealing that focusing on key diagnostic indicators can enhance model interpretability and assist medical professionals. Future research could explore additional feature selection methods and classifiers to further improve the robustness and accuracy of breast cancer classification models. The findings underscore the Random Forest model as a highly effective tool for breast cancer diagnosis, supporting its integration into clinical workflows for improved patient outcomes.
Keywords
Breast Cancer, Machine Learning, Logistic Regression, Random Forest, Support Vector Machine
References
[1] Al-Masni, M. A., Al-Azawi, R. A., & Al-Qerem, A. H. (2015). Classification of breast cancer data using artificial neural network. International Journal of Computer Science and Information Security, 13(7), 1-5.
[2] Azar A.T, El-Metwally S.M. (2012). Decision tree classifiers for automated medical diagnosis. Neural Comput Appl. 2012; 23(7–8):2387–403.
[3] Azar A.T & El-Said S.A. (2013). Performance analysis of support vector machines classifiers in breast cancer mammography recognition. Neural Comput Appl. 2013; 24(5):1163–77.
[4] Chaurasia V, Pal S, Tiwari B. (2018). Prediction of benign and malignant breast cancer using data mining techniques. J Algorithms Comput Technol. 2018; 12(2):119–26.
[5] Chen, L., Wu, M., Zhang, Z., & Wang, Y. (2018). Comparative study of breast cancer classification based on machine learning algorithms. International Journal of Hybrid Information Technology, 11(4), 333-340.
[6] Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.
[7] Hasan M.K, Islam M.M, & Hashem M.M. (2016). Mathematical model development to detect breast cancer using multi gene genetic programming. In: Proc. 5th International Conference on Informatics, Electronics and Vision (ICIEV), Dhaka, 2016, pp. 574–579.
[8] Hosmer, D.W., Lemeshow, S., & Sturdivant, R.X. (2013). Applied logistic regression. Wiley.
[9] Karabatak, M., & Ince, M.C. (2011). An expert system for detection of breast cancer based on association rules and neural network. Expert Systems with Applications, 38(7), 9010-9016.
[10] Kleinbaum, D.G., & Klein, M. (2010). Logistic Regression: A self-learning text. Springer.
[11] Mishra, N., Prakash, O., & Sinha, A. (2020). Feature selection and classification of breast cancer data using logistic regression. Journal of King Saud University - Computer and Information Sciences, 32(6), 731-736.
[12] Senapati M.R, Mohanty A.K, Dash S, & Dash P.K. (2013). Local linear wavelet neural network for breast cancer recognition. Neural Comput Appl. 2013; 22(1):125–31.
[13] Senapati M.R, Panda G, & Dash P.K. (2014). Hybrid approach using KPSO and RLS for RBFNN design for breast cancer detection. Neural Comput Appl. 2014; 24(3–4):745–53.
How to cite this paper
@article{1711021,
author = {Durowade Adeyemi Nathaniel, Ajayi Oyewole, Ayobami Samuel O, Sanni Bello, Saheed Ajibade},
title = {Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {4},
pages = {45-52},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1711021.pdf},
abstract = {This study conducts a comparative analysis of three prominent machine learning models?Logistic Regression, Random Forest, and Support Vector Machine (SVM),?for the classification of breast cancer. The data used for this study is collected from the records of incoming patients of breast cancer in Ekiti State University Teaching Hospital, Ekiti State, Nigeria. The analysis emphasizes the significance of certain features, particularly the worst measurements of texture, radius, and area, in distinguishing between benign and malignant tumors. Key features such as `area worst`, `radius worst`, and `texture worst` demonstrate high phi values, indicating their strong association with the class labels. This suggests that extreme values of these measurements are crucial in identifying malignancies. Among the evaluated models, the Random Forest model exhibits the highest accuracy, as validated by cross-validation techniques. The selection of `mtry = 2` as the optimal parameter underscores the importance of choosing the appropriate number of features at each split to maximize model performance. The model's reliability is further confirmed by confusion matrices, which show high sensitivity and specificity, critical for minimizing false negatives and positives in medical diagnoses. This study highlights the importance of feature importance analysis in medical data classification, revealing that focusing on key diagnostic indicators can enhance model interpretability and assist medical professionals. Future research could explore additional feature selection methods and classifiers to further improve the robustness and accuracy of breast cancer classification models. The findings underscore the Random Forest model as a highly effective tool for breast cancer diagnosis, supporting its integration into clinical workflows for improved patient outcomes.},
keywords = {Breast Cancer, Machine Learning, Logistic Regression, Random Forest, Support Vector Machine},
month = {October},
doi = {https://doi.org/10.64388/IREV9I4-1711021-6791}
}