Home / Current Issue / Paper 1723814
Modeling Enterprise Behavior for Anomaly Discovery in Security Event Data
Subject area: Science,Engineering and Technology · Area of research: Security Event Data
Abstract
This research presents and assesses a reproducible framework for detecting behavioral anomalies in enterprise security event logs through unsupervised machine learning. Because access to authentic security telemetry is often limited, a synthetic enterprise log dataset was constructed to represent authentication events, user activity patterns, session behavior, and simulated adversarial activities. The resulting dataset consists of 6,244 security events, including 5,874 normal records and 370 injected attack records covering credential stuffing, privilege misuse, abnormal session duration, and lateral movement. Behavioral and temporal characteristics were derived through rolling activity metrics, statistical deviation measures, user-specific behavioral indicators, and a composite risk score. Five unsupervised anomaly-detection techniques—Isolation Forest, Local Outlier Factor (LOF), One-Class SVM, Autoencoder, and Gaussian Mixture Model (GMM)—were developed using Python, scikit-learn, and PyTorch. Attack labels were intentionally withheld during model training and subsequently used only for post hoc performance assessment based on precision, recall, F1-score, accuracy, ROC-AUC, and PR-AUC. The evaluation indicates that Isolation Forest obtained the strongest fixed-threshold F1-score (0.4422) and PR-AUC (0.4608), whereas GMM recorded the highest ROC-AUC (0.8227). LOF demonstrated the lowest overall effectiveness, while One-Class SVM and the Autoencoder showed moderate performance across the evaluated measures. These results indicate that unsupervised learning techniques can identify useful anomaly patterns within engineered security-event features; however, their effectiveness is influenced considerably by feature representation, underlying algorithmic assumptions, and the selected evaluation measures. The proposed framework offers a controlled and reproducible basis for examining behavioral anomaly detection in security-monitoring environments where sufficiently labeled datasets are unavailable or limited.
References
[1] Kent K, Souppaya M. Guide to computer security log management. NIST. 2006.
[2] Sommer R, Paxson V. Outside the closed world: On using machine learning for network intrusion detection. IEEE S&P. 2010.
[3] Chandola V, et al. Anomaly detection: A survey. ACM Computing Surveys. 2009.
[4] Liu FT, et al. Isolation forest. In: ICDM. 2008.
[5] Breunig M, et al. LOF: Identifying density-based local outliers. In: SIGMOD. 2000.
[6] Schölkopf B, et al. Support vector method for novelty detection. Neural Computation. 2001.
[7] Du M, et al. DeepLog. In: ACM CCS. 2017.
[8] An J, Cho S. Autoencoder-based anomaly detection. 2015.
[9] Garcia S, et al. An empirical comparison of botnet detection methods. 2014.
[10] Hussein SA, Répás SR. Anomaly detection in log files based on machine learning techniques. 2024.
[11] Wibowo RA, Syalsabilla AF, Rozy AF. Identification of cyber attacks based on network anomaly factor in Indonesia using fuzzy Gaussian mixture model with intervention analysis. 2026.
[12] Kumar Y, Saraswat N, Gupta PK, Mathur M. Comparative evaluation of anomaly detection algorithms for network intrusion detection. In: Data Processing and Networking (ICDPN 2025). Cham: Springer; 2026. (Lecture Notes in Networks and Systems; vol. 1934).
[13] Agyemang EF. Anomaly detection using unsupervised machine learning algorithms: A simulation study. Scientific African. 2024;26:e02386.
[14] Benova L, Hudec L. Comprehensive analysis and evaluation of anomalous user activity in web server logs. Sensors. 2024;24(3):746.
[15] Gutierrez-Portela F, Almenares Mendoza F, Calderon-Benavides L. Evaluation of the performance of unsupervised learning algorithms for intrusion detection in unbalanced data environments. IEEE Access. 2024;12:190134–190157.
[16] Pennada SSP, Nayak SK, V. K. M. Insider threat detection using behavioural analysis through machine learning and deep learning techniques. International Research Journal of Multidisciplinary Technovation. 2025;7(2):74–86.
[17] Guo H, Yuan S, Wu X. LogBERT: Log anomaly detection via BERT. In: 2021 International Joint Conference on Neural Networks (IJCNN), IEEE. 2021. p. 1–8.
[18] Fei K, Zhou J, Zhou Y, Gu X, Fan H, Li B, et al. LaAeb: A comprehensive log-text analysis based approach for insider threat detection. Computers & Security. 2025;148:104126.
[19] Blanco R, Malagón P, Briongos S, Moya JM. Anomaly detection using Gaussian mixture probability model to implement intrusion detection system. In: Hybrid Artificial Intelligent Systems (HAIS 2019). Cham: Springer; 2019. p. 648–659. (Lecture Notes in Computer Science; vol. 11734).
[20] Das BC, Sartaz MS, Reza SA, Hossain A, Nasiruddin M, Bishnu KK, et al. AI-driven cybersecurity threat detection: Building resilient defense systems using predictive analytics. International Journal of Basic and Applied Sciences. 2025;14(4):33–45.
How to cite this paper
@article{1723814,
author = {Ejiaku Allen Chigozie, Onyekachi Christopher Eze, Anyanwu Prosper Chukwudi, Chinonso Valentine Nnachetam},
title = {Modeling Enterprise Behavior for Anomaly Discovery in Security Event Data},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {4},
pages = {1105-1156},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1723814.pdf},
abstract = {This research presents and assesses a reproducible framework for detecting behavioral anomalies in enterprise security event logs through unsupervised machine learning. Because access to authentic security telemetry is often limited, a synthetic enterprise log dataset was constructed to represent authentication events, user activity patterns, session behavior, and simulated adversarial activities. The resulting dataset consists of 6,244 security events, including 5,874 normal records and 370 injected attack records covering credential stuffing, privilege misuse, abnormal session duration, and lateral movement. Behavioral and temporal characteristics were derived through rolling activity metrics, statistical deviation measures, user-specific behavioral indicators, and a composite risk score. Five unsupervised anomaly-detection techniques—Isolation Forest, Local Outlier Factor (LOF), One-Class SVM, Autoencoder, and Gaussian Mixture Model (GMM)—were developed using Python, scikit-learn, and PyTorch. Attack labels were intentionally withheld during model training and subsequently used only for post hoc performance assessment based on precision, recall, F1-score, accuracy, ROC-AUC, and PR-AUC. The evaluation indicates that Isolation Forest obtained the strongest fixed-threshold F1-score (0.4422) and PR-AUC (0.4608), whereas GMM recorded the highest ROC-AUC (0.8227). LOF demonstrated the lowest overall effectiveness, while One-Class SVM and the Autoencoder showed moderate performance across the evaluated measures. These results indicate that unsupervised learning techniques can identify useful anomaly patterns within engineered security-event features; however, their effectiveness is influenced considerably by feature representation, underlying algorithmic assumptions, and the selected evaluation measures. The proposed framework offers a controlled and reproducible basis for examining behavioral anomaly detection in security-monitoring environments where sufficiently labeled datasets are unavailable or limited.},
month = {October},
}