Home / Current Issue / Paper 1718759
Behavioral Anomaly Detection in Social Security Claims Using a Longitudinal Statistical Profiling Approach with Income Variance Metrics
Subject area: Science,Engineering and Technology · Area of research: Mathematics and Statistics
DOI: 10.64388/IREV9I12-1718759
Abstract
Social security fraud imposes substantial fiscal and equity costs on pension systems worldwide, yet systematic statistical methods for its detection remain underdeveloped, particularly in resource limited institutional settings. This paper develops and validates a longitudinal behavioral profiling framework for detecting anomalous claims in a social security context. We introduce the Income Variance Ratio (IVR), a metric quantifying the proportional divergence between an individual’s documented earnings trajectory and their claimed benefit entitlement, and demonstrate its theoretical and empirical superiority over conventional scalar risk scores. Building on a formal probabilistic model of claim generation, we derive a set of structural theorems characterizing the stochastic separation between legitimate and fraudulent claim populations, and establish consistency and convergence guarantees for the proposed estimators. Empirically, we apply the framework to a longitudinal dataset of 10,000 social security claims spanning 18 months, of which 746 (7.46%) were confirmed fraudulent. All eight primary features exhibit statistically significant distributional separation (Mann-Whitney p < 10−13; Kolmogorov-Smirnov p < 10−10). The IVR achieves the largest rank-bi-serial correlation (r = 0.599) among all features. A Random Forest classifier trained on the full feature set attains a cross-validated AUC of 0.826 and a perfect confusion matrix at a threshold of 0.40 on held-out data. Critically, IVR decile analysis reveals a non-linear, threshold-crossing pattern in which the top decile concentrates 42.2% fraud prevalence against a baseline of 7.46%, corresponding to a 5.7-fold lift. The Documentation Completeness × Geographic Risk interaction surface exhibits cells with fraud rates exceeding 100% in small-sample high-risk strata. These findings provide both theoretical grounding and practical guidelines for deploying income-variance-based fraud screening in pension and social security administration.
Keywords
Social Security Fraud Detection, Income Variance Ratio, Behavioral Anomaly Scoring, Longitudinal Statistical Profiling, Random Forest Classification, Mann-Whitney Test, Principal Component Analysis, Unsupervised Anomaly Detection.
References
[1] Albrecht, W. S., Albrecht, C. O., Albrecht, C. C., and Zimbelman, M. F. (2012). Fraud Examination. South-Western Cengage Learning, Mason, OH, 4th edition.
[2] Baesens, B., Van Vlasselaer, V., and Verbeke, W. (2015). Fraud Analytics Using Descriptive, Predictive, and Social Network Techniques: A Guide to Data Science for Fraud Detection. Wiley, Hoboken, NJ.
[3] Bamber, D. (1975). The area above the ordinal dominance graph and the area below the receiver operating characteristic graph. Journal of Mathematical Psychology, 12(4):387–415.
[4] Bauder, R. A. and Khoshgoftaar, T. M. (2017). Medicare fraud detection using machine learning methods. pages 858–865.
[5] Benford, F. (1938). The law of anomalous numbers. Proceedings of the American Philosophical Society, 78(4):551–572.
[6] Bhattacharyya, S., Jha, S., Tharakunnel, K., and Westland, J. C. (2011). Data mining for credit card fraud: A comparative study. Decision Support Systems, 50(3):602–613.
[7] Bolton, R. J. and Hand, D. J. (2002). Statistical fraud detection: A review. Statistical Science, 17(3):235–255.
[8] Breiman, L. (2001). Random forests. Machine Learning, 45(1):5–32.
[9] Breunig, M. M., Kriegel, H.-P., Ng, R. T., and Sander, J. (2000). LOF: Identifying density- based local outliers. In Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, pages 93–104. ACM.
[10] Carcillo, F., Dal Pozzolo, A., Le Borgne, Y.- A., Caelen, O., Mazzer, Y., and Bontempi, G. (2018). SCARFF: A scalable framework for streaming credit card fraud detection with Spark. Information Fusion, 41:182–194.
[11] Chalapathy, R. and Chawla, S. (2019). Deep learning for anomaly detection: A survey. arXiv preprint arXiv:1901.03407.
[12] Chandola, V., Banerjee, A., and Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3):15:1–15:58.
[13] Chawla, N. V., Bowyer, K. W., Hall, L. O., and Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16:321–357.
[14] Chen, C., Liaw, A., and Breiman, L. (2004). Using random forest to learn imbalanced data. Number 666. Technical Report, Department of Statistics, University of California Berkeley.
[15] Cuzick, J. (1985). A Wilcoxon-type test for trend. Statistics in Medicine, 4(4):543–547.
[16] Dal Pozzolo, A., Caelen, O., Le Borgne, Y.- A., Waterschoot, S., and Bontempi, G. (2014). Learned lessons in credit card fraud detection from a practitioner perspective. Expert Systems with Applications, 41(10):4915– 4928.
[17] Davis, J. and Goadrich, M. (2006). The relationship between precision-recall and ROC curves. pages 233–240.
[18] Hanley, J. A. and McNeil, B. J. (1982). The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology, 143(1):29–36.
[19] He, H. and Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9):1263–1284.
[20] Joudaki, H., Rashidian, A., Minaei-Bidgoli, B., Mahmoodi, M., Geraili, B., Nasiri, M., and Arab, M. (2015). Using data mining to detect health care fraud and abuse: A review of literature. Global Journal of Health Science, 7(1):194–202.
[21] Kerby, D. S. (2014). The simple difference formula: An approach to teaching nonparametric correlation. Comprehensive Psychology, 3:11.IT.3.1.
[22] Kolmogorov, A. N. (1933). Sulla determinazione empirica di una legge di distribuzione (On the empirical determination of a distribution law). Giornale dell’Istituto Italiano degli Attuari, 4:83–91.
[23] Little, R. J. A. and Rubin, D. B. (2002). Statistical Analysis with Missing Data. Wiley, Hoboken, NJ, 2nd edition.
[24] Liu, F. T., Ting, K. M., and Zhou, Z.-H. (2008). Isolation forest. In Proceedings of the 8th IEEE International Conference on Data Mining (ICDM), pages 413–422. IEEE.
[25] Mann, H. B. and Whitney, D. R. (1947). On a test of whether one of two random variables is stochastically larger than the other. The Annals of Mathematical Statistics, 18(1):50– 60.
[26] Ngai, E. W. T., Hu, Y., Wong, Y. H., Chen, Y., and Sun, X. (2011). The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature. Decision Support Systems, 50(3):559–569.
[27] Nigrini, M. J. (1999). I’ve Got Your Number: How a Mathematical Phenomenon Can Help CPAs Uncover Fraud and Other Irregularities, volume 187. Journal of Accountancy, May 1999.
[28] Phua, C., Lee, V. C. S., Smith-Miles, K., and Gayler, R. W. (2010). A comprehensive survey of data mining-based fraud detection research. arXiv preprint arXiv:1009.6119.
[29] Pickett, K. H. S. (2011). The Essential Guide to Internal Auditing. Wiley, Chichester, UK, 2nd edition.
[30] Pudney, S., Hancock, R., and Sutherland, H. (2004). Simulating the reform of means-tested benefits with endogenous take-up and claim costs. University of Essex, Institute for Social and Economic Research (ISER), Colchester
[31] Smirnov, N. V. (1948). Table for estimating the goodness of fit of empirical distributions. The Annals of Mathematical Statistics, 19(2):279–281.
[32] Van de Walle, D., Nead, K and World Bank. (1995). Public spending and the poor: Theory and evidence. World Bank Publications.
[33] Wand, M. P. and Jones, M. C. (1994). Kernel Smoothing. Chapman and Hall, London.
[34] World Health Organization (2019). Health care fraud. Technical report, World Health Organization.
[35] Yaniv, G. (1997). Welfare fraud and welfare stigma. Journal of Economic Psychology, 18(4):435–451.
How to cite this paper
@article{1718759,
author = {Ometan S. Olokor, Kayoh O. Clinton, Ekuma-Okereke Eyinnaya, Okedoye M. Akindele},
title = {Behavioral Anomaly Detection in Social Security Claims Using a Longitudinal Statistical Profiling Approach with Income Variance Metrics},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {12},
pages = {1950-1966},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718759.pdf},
abstract = {Social security fraud imposes substantial fiscal and equity costs on pension systems worldwide, yet systematic statistical methods for its detection remain underdeveloped, particularly in resource limited institutional settings. This paper develops and validates a longitudinal behavioral profiling framework for detecting anomalous claims in a social security context. We introduce the Income Variance Ratio (IVR), a metric quantifying the proportional divergence between an individual’s documented earnings trajectory and their claimed benefit entitlement, and demonstrate its theoretical and empirical superiority over conventional scalar risk scores. Building on a formal probabilistic model of claim generation, we derive a set of structural theorems characterizing the stochastic separation between legitimate and fraudulent claim populations, and establish consistency and convergence guarantees for the proposed estimators. Empirically, we apply the framework to a longitudinal dataset of 10,000 social security claims spanning 18 months, of which 746 (7.46%) were confirmed fraudulent. All eight primary features exhibit statistically significant distributional separation (Mann-Whitney p < 10−13; Kolmogorov-Smirnov p < 10−10). The IVR achieves the largest rank-bi-serial correlation (r = 0.599) among all features. A Random Forest classifier trained on the full feature set attains a cross-validated AUC of 0.826 and a perfect confusion matrix at a threshold of 0.40 on held-out data. Critically, IVR decile analysis reveals a non-linear, threshold-crossing pattern in which the top decile concentrates 42.2% fraud prevalence against a baseline of 7.46%, corresponding to a 5.7-fold lift. The Documentation Completeness × Geographic Risk interaction surface exhibits cells with fraud rates exceeding 100% in small-sample high-risk strata. These findings provide both theoretical grounding and practical guidelines for deploying income-variance-based fraud screening in pension and social security administration.},
keywords = {Social Security Fraud Detection, Income Variance Ratio, Behavioral Anomaly Scoring, Longitudinal Statistical Profiling, Random Forest Classification, Mann-Whitney Test, Principal Component Analysis, Unsupervised Anomaly Detection.},
month = {June},
doi = {https://doi.org/10.64388/IREV9I12-1718759}
}