Home / Current Issue / Paper 1706978
Comparative Analysis of Machine Learning Algorithms for Diabetes Prediction
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
Abstract
Diabetes is regarded as one of the most chronic metabolic diseases if left unchecked. Diabetes is a worldwide chronic health issue. Today, approximately 400 million people are living with diabetes. A large percentage of people who are living with diabetes are unaware of their condition until it becomes chronic. Diabetes, also known as Diabetes Mellitus, is an increasingly prevalent chronic disease which affects the body?s ability to metabolize glucose. With the growing rate of diabetes cases, it has become important to take a deeper look into solutions and ways to better handle the situation. This paper presents a predictive approach to diabetes, through diabetes prediction using machine learning, a process that will allow for better treatment and preventive healthcare. Machine learning in diabetes prediction is important because there is a vast pool of available data on diabetes both through research and years of clinical studies. This data can be processed and fed into machine learning models to highlight meaningful relationships and patterns within patients? data. However, this has been hampered by the difficult task of choosing the best machine learning algorithm. A challenge that can be solved by carrying out a comparative study using different evaluation metrics to ascertain which algorithm produces the most optimal results. This paper represents the result and analysis regarding detecting a person?s diabetic state from various machine learning models based on key attributes such as age, gender, glucose level and insulin level. The model proposed was achieved by collating diabetes data from Kaggle and prepossessed to remove abnormalities and irrelevant attributes after which it was divided into test and training data. The machine learning algorithms chosen for this study were SVM, logistic regression, decision tree, random forest classifier and K-Neighbors classifier. The best performing model was random forest with an accuracy of 95%. This paper contributes to the diagnosis and prediction of diabetes through the application of machine learning in predicting patients who are likely to live with diabetes.
Keywords
Diabetes, Decision Tree, Dataset, Attributes, Machine Learning, SVM, K-Neighbors, Random Forest Algorithms.
References
[1] Bohr, Adam, and Kaveh Memarzadeh. “The Rise of Artificial Intelligence in Healthcare Applications.” Artificial Intelligence in Healthcare, June 2020, pp. 25–60, https://doi.org/10.1016/B978-0-12-818438-7.00002-2.
[2] Davenport, Thomas, and Ravi Kalakota. “The Potential for Artificial Intelligence in Healthcare.” Future Healthcare Journal, vol. 6, no. 2, June 2019, pp. 94–98, https://doi.org/10.7861/futurehosp.6-2-94.
[3] Girasa, Rosario. “AI as a Disruptive Technology.” Springer Link, edited by Rosario Girasa, Springer International Publishing, 2020, pp. 3–21, link.springer.com/chapter/10.1007%2F978-3-030-35975-1_1.
[4] Whelan, Rob. “Understanding Data Science, Artificial Intelligence, and Machine Learning.” 2nd Watch, 13 Jan. 2021, www.2ndwatch.com/blog/understanding-basics-data-science- artificial-intelligence-machine-learning/.
[5] Hu, F. B. “Globalization of Diabetes: The Role of Diet, Lifestyle, and Genes.” Diabetes Care, vol. 34, no. 6, May 2011, pp. 1249–57, https://doi.org/10.2337/dc11-0442.
[6] Sivarajah, Uthayasankar, et al. “Critical Analysis of Big Data Challenges and Analytical Methods.” Journal of Business Research, vol. 70, Jan. 2017, pp. 263–86, https://doi.org/10.1016/j.jbusres.2016.08.001.
[7] Kazzazi, Fawz. “The Automation of Doctors and Machines: A Classification for AI in Medicine (ADAM Framework).” Future Healthcare Journal, vol. 8, no. 2, May 2021, pp.e257–62, https://doi.org/10.7861/fhj.2020-0189.
[8] Gliklich, RE, et al. IEEE Standard for an Architectural Framework for the Internet of Things (IoT). New York, Usa Ieee, 2020.
[9] Arentze, T. A. “Spatial Data Mining, Cluster and Pattern Recognition.” International Encyclopedia of Human Geography, 2009, pp. 325–31, https://doi.org/10.1016/b978-008044910-4.00524-1.
[10] Dash, Sabyasachi, et al. “Big Data in Healthcare: Management, Analysis and Future Prospects.” Journal of Big Data, vol. 6, no. 1, June 2019, https://doi.org/10.1186/s40537-019-0217-0.
[11] Koh, Hian Chye, and Gerald Tan. “Data Mining Applications in Healthcare.” Journal of Healthcare Information Management: JHIM, vol. 19, no. 2, 2005, pp. 64–72, pubmed.ncbi.nlm.nih.gov/15869215/
[12] Palanisamy, Venketesh, and Ramkumar Thirunavukarasu. “Implications of Big Data Analytics in Developing Healthcare Frameworks – a Review.” Journal of King Saud University - Computer and Information Sciences, vol. 31, no. 4, Oct. 2019, pp. 415–25, https://doi.org/10.1016/j.jksuci.2017.12.007.
[13] Lai, Hang, et al. “Predictive Models for Diabetes Mellitus Using Machine Learning Techniques.” BMC Endocrine Disorders, vol. 19, no. 1, Oct. 2019, https://doi.org/10.1186/s12902-019-0436-6.
[14] Soni, Ankit Narendrakumar. “Diabetes Mellitus Prediction Using Ensemble Machine Learning Techniques.” SSRN Electronic Journal, vol. 9, no. 9, 2020, https://doi.org/10.2139/ssrn.3642877.
[15] ukani. “Diabetes Data Set.” Kaggle.com, 2020, www.kaggle.com/vikasukani/diabetes-data-set.
[16] Kaushik, Saurav. “An Introduction to Clustering & Different Methods of Clustering.” Analytics Vidhya, 11 Mar. 2019, www.analyticsvidhya.com/blog/2016/11/an-introduction-to- clustering-and-different-methods-of-clustering/.
[17] Bandyopadhyay, Sanghamitra. Unsupervised Classification : Similarity Measures, Classical and Metaheuristic Approaches, and Applications. Springer, 2013.
[18] Seliya, Naeem, et al. “A Study on the Relationships of Classifier Performance Metrics.” 2009 21st IEEE International Conference on Tools with Artificial Intelligence, Nov. 2009, https://doi.org/10.1109/ictai.2009.25.
[19] Alberg, Anthony J., et al. “The Use of ‘Overall Accuracy’ to Evaluate the Validity of Screening or Diagnostic Tests.” Journal of General Internal Medicine, vol. 19, no. 5, May 2004, pp. 460–65, https://doi.org/10.1111/j.1525-1497.2004.30091.x.
[20] Armah, Gabriel Kofi, et al. “A Deep Analysis of the Precision Formula for Imbalanced Class Distribution.” International Journal of Machine Learning and Computing, vol. 4, no. 5, 2014, pp. 417–22, https://doi.org/10.7763/ijmlc.2014.v4.447.
[21] Baratloo, Alireza, et al. “Part 1: Simple Definition and Calculation of Accuracy, Sensitivity and Specificity.” Emergency, vol. 3, no. 2, 2015, pp. 48–49, www.ncbi.nlm.nih.gov/pmc/articles/PMC4614595/.
[22] KLM`a¯°±²³³ ³³#³!›œ¨©ª›hÖ^hÖ^hÖ^hÖ^CJOJQJaJ hÖ^hÖ^@ˆþÿCJOJQJaJ"hÖ^hÖ^5�CJOJQJ\�aJ"hÖ^hÖ^5�6�CJOJQJaJh^~ h˜gCJOJQJaJh^~ CJOJQJaJh^~ h=°CJOJQJaJhñM³he6�OJPJQJheOJPJQJhÖ^hñM³OJPJQJh^~ h^~ OJQJh^~ h^~ CJOJQJhÖ^hÖ^CJ(OJQJLM°±²³!œ©ªòåл©©Ÿ—uk $"$ &F [Some characters in this reference could not be displayed correctly — please refer to the published PDF for the full reference.]
[23] ƃ„ƒ„Êýdð¤1$7$8$^„ƒ`„Êýa$gdÖ^m$$a$gdÖ^ $„ñÿ„ dð¤]„ñÿ^„ gd^~ $„ñÿ„ dð¤]„ñÿ^„ a$gd^~ $„`„öÿdë¤^„``„öÿa$gdÖ^$dð¤a$gd^~ $dݤa$gd˜g ªÓ $"$ &F [Some characters in this reference could not be displayed correctly — please refer to the published PDF for the full reference.]
How to cite this paper
@article{1706978,
author = {Omoshola Ogunduboye },
title = {Comparative Analysis of Machine Learning Algorithms for Diabetes Prediction},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {8},
number = {7},
pages = {527-536},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1706978.pdf},
abstract = {Diabetes is regarded as one of the most chronic metabolic diseases if left unchecked. Diabetes is a worldwide chronic health issue. Today, approximately 400 million people are living with diabetes. A large percentage of people who are living with diabetes are unaware of their condition until it becomes chronic. Diabetes, also known as Diabetes Mellitus, is an increasingly prevalent chronic disease which affects the body?s ability to metabolize glucose. With the growing rate of diabetes cases, it has become important to take a deeper look into solutions and ways to better handle the situation. This paper presents a predictive approach to diabetes, through diabetes prediction using machine learning, a process that will allow for better treatment and preventive healthcare. Machine learning in diabetes prediction is important because there is a vast pool of available data on diabetes both through research and years of clinical studies. This data can be processed and fed into machine learning models to highlight meaningful relationships and patterns within patients? data. However, this has been hampered by the difficult task of choosing the best machine learning algorithm. A challenge that can be solved by carrying out a comparative study using different evaluation metrics to ascertain which algorithm produces the most optimal results. This paper represents the result and analysis regarding detecting a person?s diabetic state from various machine learning models based on key attributes such as age, gender, glucose level and insulin level. The model proposed was achieved by collating diabetes data from Kaggle and prepossessed to remove abnormalities and irrelevant attributes after which it was divided into test and training data. The machine learning algorithms chosen for this study were SVM, logistic regression, decision tree, random forest classifier and K-Neighbors classifier. The best performing model was random forest with an accuracy of 95%. This paper contributes to the diagnosis and prediction of diabetes through the application of machine learning in predicting patients who are likely to live with diabetes.},
keywords = {Diabetes, Decision Tree, Dataset, Attributes, Machine Learning, SVM, K-Neighbors, Random Forest Algorithms.},
month = {January},
}