International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1716773

1716773 Vol 9 · Issue 10 Download Paper

Customer Churning Analysis Using Machine Learning Algorithms

Ashish Pandey Anurag Pal Tajendra Riya Prof. (Dr.) Sanjay Pachauri

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

DOI: https://doi.org/10.64388/IREV9I10-1716773

Abstract

In the contemporary subscription-based economy, customer retention is structurally paramount to the long-term viability and profitability of telecommunications enterprises. The financial asymmetry between Customer Acquisition Cost (CAC) and Customer Retention Cost (CRC) mandates the development of highly accurate, proactive customer churn prediction systems. This comprehensive research paper presents an in-depth empirical analysis of machine learning algorithms deployed to forecast customer defection. We rigorously evaluate a spectrum of predictive models, progressing from traditional linear classifiers (Logistic Regression) to complex, non-linear distance-based models (Support Vector Machines), and state-of-the-art ensemble architectures (Random Forest and eXtreme Gradient Boosting). A critical challenge in churn analytics—the severe class imbalance inherent in real-world datasets—is addressed through the application of the Synthetic Minority Over-sampling Technique (SMOTE) combined with rigorous cross-validation strategies. Utilizing a robust telecommunications dataset, our extensive feature engineering and hyperparameter optimization reveal that tree-based ensemble methods, particularly XGBoost, significantly outperform baseline models. XGBoost achieved a superior Area Under the Receiver Operating Characteristic Curve (ROC-AUC) of 0.88, optimizing the critical trade-off between Precision and Recall. Furthermore, the integration of SHapley Additive exPlanations (SHAP) provides a granular, interpretable analysis of feature importance, identifying contract duration, customer tenure, and specific service combinations as the primary drivers of attrition. The findings equip business stakeholders with interpretable, high-fidelity predictive intelligence required to operationalize targeted, cost-effective customer retention interventions.

Keywords

Customer Churn, Machine Learning, eXtreme Gradient Boosting (XGBoost), Synthetic Minority Over-sampling Technique (SMOTE), Predictive Analytics, SHAP Values, Class Imbalance, Telecommunications.

References

[1] Burez, J., & Van den Poel, D. (2009). Handling class imbalance in customer churn prediction. Expert Systems with Applications, 36(3), 4626-4636. https://doi.org/10.1016/j.eswa.2008.05.027

[2] Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321-357. https://doi.org/10.1613/jair.953

[3] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). ACM.

[4] Coussement, K., Lessmann, S., & Verstraeten, G. (2020). A comparative analysis of data preparation algorithms for customer churn prediction: A case study in the telecommunication industry. Decision Support Systems, 95, 27-36.

[5] Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd ed.). Springer.

[6] Jain, H., Khunteta, A., & Srivastava, S. (2021). Churn prediction in telecommunication using machine learning algorithms. International Journal of Advanced Computer Science and Applications, 11(12), 1-8.

[7] Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4765-4774).

[8] Neslin, S. A., Gupta, S., Kamakura, W., Lu, J., & Mason, C. H. (2006). Defection detection: Measuring and understanding the predictive accuracy of customer churn models. Journal of Marketing Research, 43(2), 204-211.

[9] Provost, F., & Fawcett, T. (2013). Data Science for Business: What you need to know about data mining and data-analytic thinking. O'Reilly Media, Inc.

[10] Verbeke, W., Martens, D., Mues, C., & Baesens, B. (2011). Building comprehensible customer churn prediction models with advanced rule induction techniques. Expert Systems with Applications, 38(3), 2354-2364.

How to cite this paper

Ashish Pandey, Anurag Pal, Tajendra Riya, Prof. (Dr.) Sanjay Pachauri "Customer Churning Analysis Using Machine Learning Algorithms" Iconic Research And Engineering Journals Volume 9 Issue 10 2026 Page 2643-2649 https://doi.org/10.64388/IREV9I10-1716773
Ashish Pandey, Anurag Pal, Tajendra Riya, Prof. (Dr.) Sanjay Pachauri "Customer Churning Analysis Using Machine Learning Algorithms" Iconic Research And Engineering Journals, vol. 9, no. 10, Apr. 2026, doi: https://doi.org/10.64388/IREV9I10-1716773
Ashish Pandey, Anurag Pal, Tajendra Riya, Prof. (Dr.) Sanjay Pachauri (2026). Customer Churning Analysis Using Machine Learning Algorithms. Iconic Research And Engineering Journals, 9(10). doi: https://doi.org/10.64388/IREV9I10-1716773
Ashish Pandey, Anurag Pal, Tajendra Riya, Prof. (Dr.) Sanjay Pachauri "Customer Churning Analysis Using Machine Learning Algorithms" Iconic Research And Engineering Journals, vol. 9, no. 10, Apr. 2026. Crossref, https://doi.org/10.64388/IREV9I10-1716773
@article{1716773,
      author = {Ashish Pandey, Anurag Pal, Tajendra Riya, Prof. (Dr.) Sanjay Pachauri},
      title = {Customer Churning Analysis Using Machine Learning Algorithms},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {10},
      pages = {2643-2649},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1716773.pdf},
      abstract = {In the contemporary subscription-based economy, customer retention is structurally paramount to the long-term viability and profitability of telecommunications enterprises. The financial asymmetry between Customer Acquisition Cost (CAC) and Customer Retention Cost (CRC) mandates the development of highly accurate, proactive customer churn prediction systems. This comprehensive research paper presents an in-depth empirical analysis of machine learning algorithms deployed to forecast customer defection. We rigorously evaluate a spectrum of predictive models, progressing from traditional linear classifiers (Logistic Regression) to complex, non-linear distance-based models (Support Vector Machines), and state-of-the-art ensemble architectures (Random Forest and eXtreme Gradient Boosting). A critical challenge in churn analytics—the severe class imbalance inherent in real-world datasets—is addressed through the application of the Synthetic Minority Over-sampling Technique (SMOTE) combined with rigorous cross-validation strategies. Utilizing a robust telecommunications dataset, our extensive feature engineering and hyperparameter optimization reveal that tree-based ensemble methods, particularly XGBoost, significantly outperform baseline models. XGBoost achieved a superior Area Under the Receiver Operating Characteristic Curve (ROC-AUC) of 0.88, optimizing the critical trade-off between Precision and Recall. Furthermore, the integration of SHapley Additive exPlanations (SHAP) provides a granular, interpretable analysis of feature importance, identifying contract duration, customer tenure, and specific service combinations as the primary drivers of attrition. The findings equip business stakeholders with interpretable, high-fidelity predictive intelligence required to operationalize targeted, cost-effective customer retention interventions.},
      keywords = {Customer Churn, Machine Learning, eXtreme Gradient Boosting (XGBoost), Synthetic Minority Over-sampling Technique (SMOTE), Predictive Analytics, SHAP Values, Class Imbalance, Telecommunications.},
      month = {April},
      doi = {https://doi.org/10.64388/IREV9I10-1716773}
  }