Home / Current Issue / Paper 1722663
AI-Driven Customer Segmentation and Loyalty Prediction for Personalised Marketing: An Interpretable Machine-Learning Analysis of a Small, High-Collinearity Customer Dataset
Subject area: Science,Engineering and Technology · Area of research: Data Analysis
Abstract
Companies of every size now collect substantial volumes of customer data, but converting that data into segments, forecasts, and explainable decisions remains difficult, particularly for student researchers and small organisations without access to large, professionally curated datasets. This study builds and transparently evaluates a complete customer-analytics pipeline on a dataset of 238 individual customers, covering data cleaning, feature engineering, exploratory analysis, K-Means segmentation, engagement-tier prediction, SHAP-based explainability, and a rule-based personalised marketing framework. K-Means clustering (k=4, validated using the Elbow and Silhouette methods) produced four interpretable segments: High-Value Loyal, Established, Growing/Potential, and New/Low-Engagement. Four classifiers — Logistic Regression, Decision Tree, Random Forest, and Gradient Boosting — were trained on demographic predictors alone (age, annual income, region) to forecast engagement tier, achieving test accuracies between 96.7% and 100%. Diagnostic analysis traces this near-perfect performance to severe multicollinearity in the dataset (Variance Inflation Factors of 37–197; a single principal component explaining 98.8% of variance) rather than to genuine predictive strength, a conclusion corroborated by SHAP, which identifies annual income as the dominant predictor, age as a weaker secondary predictor, and region as largely irrelevant. Rather than overstating predictive novelty, the study advances a narrower and more defensible claim: that a segmentation–prediction–explainability framework already validated on large industrial datasets can be meaningfully applied to, and honestly evaluated on, the small, highly collinear data typical of student and SME research. The pipeline is operationalised as an interactive dashboard application.
Keywords
customer segmentation, explainable AI, K-Means clustering, multicollinearity, personalised marketing
How to cite this paper
@article{1722663,
author = {P. T. Manasa Visakai, Sadhana Venkatraghavan},
title = {AI-Driven Customer Segmentation and Loyalty Prediction for Personalised Marketing: An Interpretable Machine-Learning Analysis of a Small, High-Collinearity Customer Dataset},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {2},
pages = {3446-3460},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722663.pdf},
abstract = {Companies of every size now collect substantial volumes of customer data, but converting that data into segments, forecasts, and explainable decisions remains difficult, particularly for student researchers and small organisations without access to large, professionally curated datasets. This study builds and transparently evaluates a complete customer-analytics pipeline on a dataset of 238 individual customers, covering data cleaning, feature engineering, exploratory analysis, K-Means segmentation, engagement-tier prediction, SHAP-based explainability, and a rule-based personalised marketing framework. K-Means clustering (k=4, validated using the Elbow and Silhouette methods) produced four interpretable segments: High-Value Loyal, Established, Growing/Potential, and New/Low-Engagement. Four classifiers — Logistic Regression, Decision Tree, Random Forest, and Gradient Boosting — were trained on demographic predictors alone (age, annual income, region) to forecast engagement tier, achieving test accuracies between 96.7% and 100%. Diagnostic analysis traces this near-perfect performance to severe multicollinearity in the dataset (Variance Inflation Factors of 37–197; a single principal component explaining 98.8% of variance) rather than to genuine predictive strength, a conclusion corroborated by SHAP, which identifies annual income as the dominant predictor, age as a weaker secondary predictor, and region as largely irrelevant. Rather than overstating predictive novelty, the study advances a narrower and more defensible claim: that a segmentation–prediction–explainability framework already validated on large industrial datasets can be meaningfully applied to, and honestly evaluated on, the small, highly collinear data typical of student and SME research. The pipeline is operationalised as an interactive dashboard application.},
keywords = {customer segmentation, explainable AI, K-Means clustering, multicollinearity, personalised marketing},
month = {August},
}