Home / Current Issue / Paper 1722663
AI-Driven Customer Segmentation and Loyalty Prediction for Personalised Marketing: An Interpretable Machine-Learning Analysis of a Small, High-Collinearity Customer Dataset
Subject area: Science,Engineering and Technology · Area of research: Data Analysis
DOI: https://doi.org/10.64388/IREV10I2-1722663
Abstract
Companies of every size now collect substantial volumes of customer data, but converting that data into segments, forecasts, and explainable decisions remains difficult, particularly for student researchers and small organisations without access to large, professionally curated datasets. This study builds and transparently evaluates a complete customer-analytics pipeline on a dataset of 238 individual customers, covering data cleaning, feature engineering, exploratory analysis, K-Means segmentation, engagement-tier prediction, SHAP-based explainability, and a rule-based personalised marketing framework. K-Means clustering (k=4, validated using the Elbow and Silhouette methods) produced four interpretable segments: High-Value Loyal, Established, Growing/Potential, and New/Low-Engagement. Four classifiers — Logistic Regression, Decision Tree, Random Forest, and Gradient Boosting — were trained on demographic predictors alone (age, annual income, region) to forecast engagement tier, achieving test accuracies between 96.7% and 100%. Diagnostic analysis traces this near-perfect performance to severe multicollinearity in the dataset (Variance Inflation Factors of 37–197; a single principal component explaining 98.8% of variance) rather than to genuine predictive strength, a conclusion corroborated by SHAP, which identifies annual income as the dominant predictor, age as a weaker secondary predictor, and region as largely irrelevant. Rather than overstating predictive novelty, the study advances a narrower and more defensible claim: that a segmentation–prediction–explainability framework already validated on large industrial datasets can be meaningfully applied to, and honestly evaluated on, the small, highly collinear data typical of student and SME research. The pipeline is operationalised as an interactive dashboard application.
Keywords
customer segmentation, explainable AI, K-Means clustering, multicollinearity, personalised marketing
References
[1] Brodie, R. J., Hollebeek, L. D., Jurić, B., & Ilić, A. (2011). Customer engagement: Conceptual domain, fundamental propositions, and implications for research. Journal of Service Research, 14(3), 252 -271. https://
[2] El Attar, A., & El-Hajj, M. (2026). Explainable AI-driven customer churn prediction: A multi-model ensemble approach with SHAP-based feature analysis. Frontiers in Artificial Intelligence. https://
[3] Hollebeek, L. D., Glynn, M. S., & Brodie, R. J. (2014). Consumer brand engagement in social media: Conceptualization, scale development and validation. Journal of Interactive Marketing, 28(2), 149-165. https://
[4] Hollebeek, L. D., et al. (2023). Hallmarks and potential pitfalls of customer- and consumer engagement scales: A systematic review. Psychology & Marketing. https://
[5] Lalitha, S., Reddy, M. U., Hussain, G. N., & Reddy, N. V. L. (2025). RFM analysis using K-means algorithm for customer segmentation. In Lecture Notes in Electrical Engineering (Springer). https:// 5_12
[6] Lewaaelhamd, I. (2024). Customer segmentation using machine learning model: An application of RFM analysis. Journal of Data Science and Intelligent Systems. https:// 293
[7] Martin, K. D., Borah, A., & Palmatier, R. W. (2017). Data privacy: Effects on customer and firm performance. Journal of Marketing, 81(1), 36 -58. https://
[8] The Implementation of RFM Analysis to Customer Profiling Using K -Means Clustering. (2023-24). Mathematical Modelling of Engineering Problems (IIETA). https://
[9] A Systematic Literature Review on Explainable AI for Churn Prediction, Customer Segmentation and Retention. (2025- 26). International Journal of Data Science and Analytics (Springer). https://
[10] A Data-Driven Approach with Explainable AI for Customer Churn Prediction in Telecommunications. (2025). ScienceDirect (Elsevier).
[11] A Systematic Review of Artificial Intelligence-Based Customer Analytics in Personalized Digital Marketing (2019-2026). (2026). American Journal of Data Science and Analytics.
[12] The Role of Artificial Intelligence in Customer Engagement and Social Media Marketing — Implications for Tourism and Hospitality. (2025). Journal of Theoretical and Applied Electronic Commerce Research (MDPI).
[13] A Systematic Review on AI in Marketing: Personalization of Advertising Campaigns and Impact on Customer Satisfaction. (2025). Springer conference volume. https:// 6_11
[14] AI-Driven Personalization and Customer Engagement: A PRISMA-Guided Systematic Literature and Bibliometric Review. (2026). Frontiers in Artificial Intelligence.
[15] Data Security and Privacy Concerns of AI- Driven Marketing. (2024). Cogent Business & Management (Taylor & Francis). https:// 43
[16] Product Collaborative Filtering Based Recommendation Systems for Large-Scale E- commerce. (2025). ScienceDirect (Elsevier, open access).
[17] Artificial Intelligence and Recommender Systems in E-commerce: Trends and Research Agenda. (2024). ScienceDirect (Elsevier, open access).
[18] A Personalized Product Recommendation Model in E-commerce Based on Retrieval Strategy. (2024). ScienceDirect (Elsevier, open access).
[19] Recommendation System for E-Commerce Using Collaborative Filtering. (2024). Journal Européen des Systèmes Automatisés (IIETA). https://
[20] Enhancing Customer Retention in Online Retail Through Churn Prediction: A Hybrid RFM, K-means, and Deep Neural Network Approach. (2025). ScienceDirect (Elsevier, open access).
How to cite this paper
@article{1722663,
author = {P. T. Manasa Visakai, Sadhana Venkatraghavan},
title = {AI-Driven Customer Segmentation and Loyalty Prediction for Personalised Marketing: An Interpretable Machine-Learning Analysis of a Small, High-Collinearity Customer Dataset},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {2},
pages = {3446-3460},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722663.pdf},
abstract = {Companies of every size now collect substantial volumes of customer data, but converting that data into segments, forecasts, and explainable decisions remains difficult, particularly for student researchers and small organisations without access to large, professionally curated datasets. This study builds and transparently evaluates a complete customer-analytics pipeline on a dataset of 238 individual customers, covering data cleaning, feature engineering, exploratory analysis, K-Means segmentation, engagement-tier prediction, SHAP-based explainability, and a rule-based personalised marketing framework. K-Means clustering (k=4, validated using the Elbow and Silhouette methods) produced four interpretable segments: High-Value Loyal, Established, Growing/Potential, and New/Low-Engagement. Four classifiers — Logistic Regression, Decision Tree, Random Forest, and Gradient Boosting — were trained on demographic predictors alone (age, annual income, region) to forecast engagement tier, achieving test accuracies between 96.7% and 100%. Diagnostic analysis traces this near-perfect performance to severe multicollinearity in the dataset (Variance Inflation Factors of 37–197; a single principal component explaining 98.8% of variance) rather than to genuine predictive strength, a conclusion corroborated by SHAP, which identifies annual income as the dominant predictor, age as a weaker secondary predictor, and region as largely irrelevant. Rather than overstating predictive novelty, the study advances a narrower and more defensible claim: that a segmentation–prediction–explainability framework already validated on large industrial datasets can be meaningfully applied to, and honestly evaluated on, the small, highly collinear data typical of student and SME research. The pipeline is operationalised as an interactive dashboard application.},
keywords = {customer segmentation, explainable AI, K-Means clustering, multicollinearity, personalised marketing},
month = {August},
doi = {https://doi.org/10.64388/IREV10I2-1722663}
}