International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1711021

1711021 Vol 9 · Issue 4 Download Paper

Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis

Durowade Adeyemi Nathaniel Ajayi Oyewole Ayobami Samuel O Sanni Bello Saheed Ajibade

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

DOI: https://doi.org/10.64388/IREV9I4-1711021-6791

Abstract

This study conducts a comparative analysis of three prominent machine learning models?Logistic Regression, Random Forest, and Support Vector Machine (SVM),?for the classification of breast cancer. The data used for this study is collected from the records of incoming patients of breast cancer in Ekiti State University Teaching Hospital, Ekiti State, Nigeria. The analysis emphasizes the significance of certain features, particularly the worst measurements of texture, radius, and area, in distinguishing between benign and malignant tumors. Key features such as `area worst`, `radius worst`, and `texture worst` demonstrate high phi values, indicating their strong association with the class labels. This suggests that extreme values of these measurements are crucial in identifying malignancies. Among the evaluated models, the Random Forest model exhibits the highest accuracy, as validated by cross-validation techniques. The selection of `mtry = 2` as the optimal parameter underscores the importance of choosing the appropriate number of features at each split to maximize model performance. The model's reliability is further confirmed by confusion matrices, which show high sensitivity and specificity, critical for minimizing false negatives and positives in medical diagnoses. This study highlights the importance of feature importance analysis in medical data classification, revealing that focusing on key diagnostic indicators can enhance model interpretability and assist medical professionals. Future research could explore additional feature selection methods and classifiers to further improve the robustness and accuracy of breast cancer classification models. The findings underscore the Random Forest model as a highly effective tool for breast cancer diagnosis, supporting its integration into clinical workflows for improved patient outcomes.

Keywords

Breast Cancer, Machine Learning, Logistic Regression, Random Forest, Support Vector Machine

References

[1] 0 (benign)

[2] 1 (malignant)

[3] 0 (benign)

[4] 107

[5] 5

[6] 1 (malignant)

[7] 0

[8] 59

[9] Table 6: Statistics

[10] CATEGORIES

[11] DETAILS

[12] Accuracy

[13] 0.9708

[14] 95% CI

[15] (0.9331, 0.9904)

[16] No Information Rate

[17] 0.6257

[18] P-Value [Acc > NIR]

[19] < 2e-16

[20] Kappa

[21] 0.9366

[22] Mcnemar's Test P-Value

[23] 0.07364

[24] Sensitivity

[25] 1

[26] Specificity

[27] 0.9219

[28] Positive Predictive Value

[29] 0.9554

[30] Negative Predictive Value

[31] 1

[32] Prevalence

[33] 0.6257

[34] Detection Rate

[35] 0.6257

[36] Detection Prevalence

[37] 0.655

[38] Balanced Accuracy

[39] 0.9609

[40] Positive Class

[41] 0

[42] Comparison of the Models: Here, the performances of three Machine Learning methods (Logistic Regression, Random Forest and Support Vector Machine) are compared in the table below:

[43] Table 7: Comparison of the Models

[44] Model

[45] Logistic Regression

[46] Random Forest

[47] SVM

[48] Accuracy

[49] 0.9531

[50] 0.9708

[51] 0.9620

[52] Kappa

[53] 0.9055

[54] 0.9366

[55] 0.9204

[56] Sensitivity

[57] 0.9811

[58] 1.0000

[59] 0.9906

[60] Specificity

[61] 0.8906

[62] 0.9219

[63] 0.8906

[64] Positive Predictive Value

[65] 0.9623

[66] 0.9554

[67] 0.9623

[68] Negative Predictive Value

[69] 0.9464

[70] 1.0000

[71] 0.9813

[72] Balanced Accuracy

[73] 0.9358

[74] 0.9609

[75] 0.9406

[76] The table compares the performance of three machine learning models Logistic Regression, Random Forest and Support Vector Machine. The key metrics show that the Random Forest has the highest accuracy of 0.9708 while the logistic regression has the lowest accuracy of 0.9501. Support Vector Machine also performs well with an accuracy of 0.9620. Also, random forest model has the highest Kappa (measure of agreement), sensitivity, specificity and balanced accuracy while the logistic regression has the lowest Kappa (measure of agreement), sensitivity, specificity and balanced accuracy.

[77] This means that random forest model is the best overall performer across all parameters used for model assessment with the highest accuracy of 97.08% and notably having a perfect sensitivity – meaning that it correctly identifies all positive cases of breast cancer.

[78] CONCLUSION

[79] The analysis demonstrates that certain features, particularly those related to the worst measurements of texture, radius, and area, have high importance in distinguishing between benign and malignant tumors. Features like area worst, radius worst, and texture worst show significant phi values, indicating a strong association with the class labels. This suggests that the extreme values of these measurements are crucial in identifying malignancies.

[80] When compared to the other models considered in this study, the Random Forest model's high accuracy, supported by cross-validation, confirms its effectiveness for this classification task. The choice of mtry = 4 as the optimal parameter highlights the importance of carefully selecting the number of features at each split to maximize model performance. The confusion matrix reinforces the model's reliability, showing high sensitivity and specificity, which are critical in medical diagnoses to minimize false negatives and positives.

[81] This study also underscores the significance of feature importance analysis in medical data classification. By identifying key features that contribute to accurate classifications, we can enhance the model's interpretability and potentially guide medical professionals in focusing on the most relevant diagnostic indicators. Future work could explore other feature selection methods and classifiers to further improve the robustness and accuracy of breast cancer classification models.

[82] RECOMMENDATIONS

[83] The features related to the worst measurements of texture, radius, and area (specifically, area worst, radius worst, and texture worst) have been identified as highly important for distinguishing between benign and malignant tumors. Ensuring that these key features are included in the breast cancer dataset and are given priority during feature selection and preprocessing stages is recommended. Further research should be conducted to understand why these features are particularly significant and how they can be measured more accurately in clinical settings.

[84] The Random Forest model has demonstrated high accuracy and robustness for the breast cancer classification task. Implementing the Random Forest model as a major ML tool for breast cancer diagnosis in the clinical workflow will help advance improvement of breast cancer care.

[85] Data scientists and Statisticians working in healthcare should not underemphasize carefully tuning the Random Forest and other ML model parameters, to ensure the best possible performance. Regular tuning and validation should be part of the model maintenance protocol.

[86] Having improved the reliability and robustness of the Random Forest model, cross-validation and other resampling techniques strengthens the case for their usage in ML and improving model performance over time. This helps in identifying any potential issues early and ensures that the model maintains its high accuracy.

[87] The importance of worst measurements of texture, radius, and area indicate that precise measurement techniques are crucial. Investing in improving the methods and tools used to measure these features and other important ones is a necessity in clinical settings. Ensuring high-quality, accurate data will enhance the model's predictions..

[88] REFERENCES

[89] Al-Masni, M. A., Al-Azawi, R. A., & Al-Qerem, A. H. (2015). Classification of breast cancer data using artificial neural network. International Journal of Computer Science and Information Security, 13(7), 1-5.

[90] Azar A.T, El-Metwally S.M. (2012). Decision tree classifiers for automated medical diagnosis. Neural Comput Appl. 2012; 23(7–8):2387–403.

[91] Azar A.T & El-Said S.A. (2013). Performance analysis of support vector machines classifiers in breast cancer mammography recognition. Neural Comput Appl. 2013; 24(5):1163–77.

[92] Chaurasia V, Pal S, Tiwari B. (2018). Prediction of benign and malignant breast cancer using data mining techniques. J Algorithms Comput Technol. 2018; 12(2):119–26.

[93] Chen, L., Wu, M., Zhang, Z., & Wang, Y. (2018). Comparative study of breast cancer classification based on machine learning algorithms. International Journal of Hybrid Information Technology, 11(4), 333-340.

[94] Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.

[95] Hasan M.K, Islam M.M, & Hashem M.M. (2016). Mathematical model development to detect breast cancer using multi gene genetic programming. In: Proc. 5th International Conference on Informatics, Electronics and Vision (ICIEV), Dhaka, 2016, pp. 574–579.

[96] Hosmer, D.W., Lemeshow, S., & Sturdivant, R.X. (2013). Applied logistic regression. Wiley.

[97] Karabatak, M., & Ince, M.C. (2011). An expert system for detection of breast cancer based on association rules and neural network. Expert Systems with Applications, 38(7), 9010-9016.

[98] Kleinbaum, D.G., & Klein, M. (2010). Logistic Regression: A self-learning text. Springer.

[99] Mishra, N., Prakash, O., & Sinha, A. (2020). Feature selection and classification of breast cancer data using logistic regression. Journal of King Saud University - Computer and Information Sciences, 32(6), 731-736.

[100] Senapati M.R, Mohanty A.K, Dash S, & Dash P.K. (2013). Local linear wavelet neural network for breast cancer recognition. Neural Comput Appl. 2013; 22(1):125–31.

[101] Senapati M.R, Panda G, & Dash P.K. (2014). Hybrid approach using KPSO and RLS for RBFNN design for breast cancer detection. Neural Comput Appl. 2014; 24(3–4):745–53.

How to cite this paper

Durowade Adeyemi Nathaniel, Ajayi Oyewole, Ayobami Samuel O, Sanni Bello, Saheed Ajibade "Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis" Iconic Research And Engineering Journals Volume 9 Issue 4 2025 Page 45-52 https://doi.org/10.64388/IREV9I4-1711021-6791
Durowade Adeyemi Nathaniel, Ajayi Oyewole, Ayobami Samuel O, Sanni Bello, Saheed Ajibade "Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis" Iconic Research And Engineering Journals, vol. 9, no. 4, Oct. 2025, doi: https://doi.org/10.64388/IREV9I4-1711021-6791
Durowade Adeyemi Nathaniel, Ajayi Oyewole, Ayobami Samuel O, Sanni Bello, Saheed Ajibade (2025). Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis. Iconic Research And Engineering Journals, 9(4). doi: https://doi.org/10.64388/IREV9I4-1711021-6791
Durowade Adeyemi Nathaniel, Ajayi Oyewole, Ayobami Samuel O, Sanni Bello, Saheed Ajibade "Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis" Iconic Research And Engineering Journals, vol. 9, no. 4, Oct. 2025. Crossref, https://doi.org/10.64388/IREV9I4-1711021-6791
@article{1711021,
      author = {Durowade Adeyemi Nathaniel, Ajayi Oyewole, Ayobami Samuel O, Sanni Bello, Saheed Ajibade},
      title = {Comparative Analysis of Machine Learning Algorithms for Predicting Breast Cancer Diagnosis},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {4},
      pages = {45-52},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1711021.pdf},
      abstract = {This study conducts a comparative analysis of three prominent machine learning models?Logistic Regression, Random Forest, and Support Vector Machine (SVM),?for the classification of breast cancer. The data used for this study is collected from the records of incoming patients of breast cancer in Ekiti State University Teaching Hospital, Ekiti State, Nigeria. The analysis emphasizes the significance of certain features, particularly the worst measurements of texture, radius, and area, in distinguishing between benign and malignant tumors. Key features such as `area worst`, `radius worst`, and `texture worst` demonstrate high phi values, indicating their strong association with the class labels. This suggests that extreme values of these measurements are crucial in identifying malignancies. Among the evaluated models, the Random Forest model exhibits the highest accuracy, as validated by cross-validation techniques. The selection of `mtry = 2` as the optimal parameter underscores the importance of choosing the appropriate number of features at each split to maximize model performance. The model's reliability is further confirmed by confusion matrices, which show high sensitivity and specificity, critical for minimizing false negatives and positives in medical diagnoses. This study highlights the importance of feature importance analysis in medical data classification, revealing that focusing on key diagnostic indicators can enhance model interpretability and assist medical professionals. Future research could explore additional feature selection methods and classifiers to further improve the robustness and accuracy of breast cancer classification models. The findings underscore the Random Forest model as a highly effective tool for breast cancer diagnosis, supporting its integration into clinical workflows for improved patient outcomes.},
      keywords = {Breast Cancer, Machine Learning, Logistic Regression, Random Forest, Support Vector Machine},
      month = {October},
      doi = {https://doi.org/10.64388/IREV9I4-1711021-6791}
  }