Home / Current Issue / Paper 1710356
Predicting Loan Defaults Using Big Data Analytics and Machine Learning
Subject area: Science,Engineering and Technology · Area of research: Big Data
Abstract
This research focuses on predicting loan defaults using big data analytics machine learning models applied to a comprehensive loan dataset. The analysis is conducted using R statistical software, enabling data-driven insights for enhanced credit risk management. Three algorithms Random Forest, XGBoost, and Na?ve Bayes are implemented to determine the most effective predictive model and identify key risk factors. The study utilized a comprehensive loan dataset sourced from Kaggle which comprised of 148,670 individual loan records, each characterized by 34 features spanning borrower demographics, financial characteristics, and loan specifications. Feature selection followed a multi-stage process designed to optimize model performance while maintaining interpretability. The balanced dataset (73,278) was partitioned using stratified random sampling to ensure representative class distribution. Model performance was assessed using multiple metrics to provide a comprehensive evaluation. XGBoost emerged as the optimal algorithm, achieving 80.5% accuracy through its sophisticated gradient-boosting framework and robust handling of class imbalance. The research establishes several key contributions to the field of credit risk modeling.
Keywords
Loan Predicting, Loan Defaults, Big Data, Machine Learning, Random Forest, XGBoost, and Na?ve Bayes
References
[1] Aslam, U., Tariq Aziz, H. I., Sohail, A. and N. K. Batcha (2019). “An empirical study on loan default prediction models,” J. Comput. Theor. Nanosci., vol. 16, no. 8, pp. 3483–3488.
[2] Su, C. W., Liu, F. Qin, M. and T. Chnag (2023). “Is a consumer loan a catalyst for confidence?,” Econ. Res. Istraživanja, vol. 36, no. 2, p. 2142260.
[3] Fatmawati, K. (2022). “Gross Domestic Product: Financing & Investment Activities and State Expenditures,” KINERJA J. Manaj. Organ. dan Ind., vol. 1, no. 1, pp. 11–18.
[4] Goodell, J. W., Kumar, S., Lim, W. M. and D. Pattnaik (2021). “Artificial intelligence and machine learning in finance: Identifying foundations, themes, and research clusters from bibliometric analysis,” J. Behav. Exp. Financ., vol. 32, p. 100577.
[5] Mhlanga, D. (2021). “Financial inclusion in emerging economies: The application of machine learning and artificial intelligence in credit risk assessment,” Int. J. Financ. Stud., vol. 9, no. 3, p. 39.
[6] Lee, I. and Y. J. Shin (2020). “Machine learning for enterprises: Applications, algorithm selection, and challenges,” Bus. Horiz., vol. 63, no. 2, pp. 157–170.
[7] Li, Y., & Zhong, H. (2012). A practical approach to credit risk assessment: How machine learning models are improving default prediction. Journal of Financial Risk Management, 4(2), 45–58.
[8] Marqués, J. M., García, V., & Sanchez, J. S. (2012). Exploring the effectiveness of statistical and machine learning models in credit scoring. Expert Systems with Applications, 39(3), 6000–6006.https://doi.org/10.1016/j.eswa.2011.11.057
[9] Dumitrescu, D., Copotoiu, F. M., & Stan, G. (2022). Why logistic regression still matters in credit scoring: Interpretability and regulatory compliance. Journal of Risk and Financial Management, 15(3), 120. https://doi.org/10.3390/jrfm15030120
[10] Levy, J. J., & O’Malley, A. J. (2020). The importance of explainability in AI: A case study in credit scoring. Artificial Intelligence in Finance, 6(2), 87–102. https://doi.org/10.1016/j.aif.2020.04.005
[11] Goodell, J. W., Kumar, S., Lim, K., & Pattnaik, D. (2021). Machine learning in finance: Applications and emerging trends. Finance Research Letters, 38, 101497. https://doi.org/10.1016/j.frl.2020.101497
[12] Mhlanga, D. (2021). Artificial intelligence in the financial sector: Challenges and opportunities. Journal of Applied Artificial Intelligence, 35(9), 659–672. https://doi.org/10.1080/08839514.2021.1885503
[13] Lee, I., & Shin, Y. J. (2020). Machine learning for enterprises: Applications in credit scoring and default prediction. Business Horizons, 63(2), 157–170. https://doi.org/10.1016/j.bushor.2019.12.004
[14] Wu, W.J. (2022) Machine Learning Approaches to Predict Loan Default. Intelligent Information Management, 14, 157-164. https://doi.org/10.4236/iim.2022.145011
[15] Efekodo K. O., 2Akinola O. S., and 3Waheed A. A. (2025). Evaluation of Machine Learning-Based Algorithm to Predicting Loan Default in Nigeria. University of Ibadan Journal of Science and Logics in ICT Research (UIJSLICTR) Vol. 13 No. 1 Jan. 2025 ISSN: 2714-3627
[16] Akinjole, A.; Shobayo, O.; Popoola, J.; Okoyeigbo, O.; Ogunleye, B. (2024). Ensemble-Based Machine Learning Algorithm for Loan Default Risk Prediction. Mathematics 2024, 12, 3423. https://doi.org/10.3390/ math12213423
[17] Isa, F., & Isa, R. (2021). Treatment of toxic asset by deposit money banks in Nigeria: A review of literature. TSU-International Journal of Accounting and Finance, 1(1), 42-50.
How to cite this paper
@article{1710356,
author = {Uchenna Emmanuel Evans-Anoruo},
title = {Predicting Loan Defaults Using Big Data Analytics and Machine Learning},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {2},
pages = {1258-1264},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1710356.pdf},
abstract = {This research focuses on predicting loan defaults using big data analytics machine learning models applied to a comprehensive loan dataset. The analysis is conducted using R statistical software, enabling data-driven insights for enhanced credit risk management. Three algorithms Random Forest, XGBoost, and Na?ve Bayes are implemented to determine the most effective predictive model and identify key risk factors. The study utilized a comprehensive loan dataset sourced from Kaggle which comprised of 148,670 individual loan records, each characterized by 34 features spanning borrower demographics, financial characteristics, and loan specifications. Feature selection followed a multi-stage process designed to optimize model performance while maintaining interpretability. The balanced dataset (73,278) was partitioned using stratified random sampling to ensure representative class distribution. Model performance was assessed using multiple metrics to provide a comprehensive evaluation. XGBoost emerged as the optimal algorithm, achieving 80.5% accuracy through its sophisticated gradient-boosting framework and robust handling of class imbalance. The research establishes several key contributions to the field of credit risk modeling. },
keywords = {Loan Predicting, Loan Defaults, Big Data, Machine Learning, Random Forest, XGBoost, and Na?ve Bayes},
month = {August},
}