International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1712447

1712447 Vol 7 · Issue 8 Download Paper

Enhanced K-Means Clustering Implementation for Web-Based Suspicious Profiles Detection

R. Prasanth Reddy Nagavelli Yogender Nath Gattu Ramya Syed Abdul Haq

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

DOI: 10.64388/IREV7I8-1712447

Abstract

Fraud detection often requires a hybrid approach combining both human expertise and AI techniques. While AI models can process large amounts of data and detect intricate patterns, human intervention is sometimes necessary to refine rules, especially when dealing with ambiguous or doubtful cases. This is particularly important when distinguishing between fraudulent and legitimate behavior, as subtle differences can exist. Another challenge is the multilingual nature of the web. Fraud can occur across different languages and cultural contexts, making it difficult to build comprehensive fraud detection models that account for all variations. In cases where linguistic resources (such as parsers or language models) are not readily available, ML tools can be adapted to learn from the data itself. However, for such models to function effectively, it is crucial to develop a rich set of features, ensuring that the data is representative and inclusive of the varied aspects of web-based fraud. This paper compares the performance of various clustering algorithms and develops an enhanced k-means algorithm utilizing the Minkowski metric. This modified algorithm, applied to an unlabeled dataset from an online dating site, effectively clusters users into authentic and suspicious categories. Supervised machine learning techniques are subsequently employed to validate the proposed model using another labeled online dating fraud dataset from the US. This methodology underscores the potential of combining traditional clustering methods with machine learning techniques to enhance anomaly detection across various domains, offering a practical tool for administrators and security experts.

Keywords

Deep Learning, Approaches Web-based Fraud, Fraud Detection, k-means algorithm

References

[1] Abdallah, A., Maarof, M. A., & Zainal, A. (2016). Fraud detection system: A survey. Journal of Network and Computer Applications, 68, 90-113.

[2] Adewole, K. S., Han, T., Wu, W., Song, H., & Sangaiah, A. K. (2020). Twitter spam account detection based on clustering and classification methods. The Journal of Supercomputing, 76, 4802-4837.

[3] Ahmed, M., Mahmood, A. N., & Hu, J. (2016). A survey of network anomaly detection techniques. Journal of Network and Computer Applications, 60, 19-31.

[4] Al Sarah, N., Rifat, F. Y., Hossain, M. S., & Narman, H. S. (2021). An efficient android malware prediction using Ensemble machine learning algorithms. Procedia Computer Science, 191, 184-191.

[5] Alom, Z., Carminati, B., & Ferrari, E. (2018). Detecting spam accounts on Twitter. In 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), IEEE, 1191-1198.

[6] Ayoobi, N., Shahriar, S., & Mukherjee, A. (2023). The looming threat of fake and llmgenerated LinkedIn profiles: Challenges and opportunities for detection and prevention. In Proceedings of the 34th ACM Conference on Hypertext and Social Media, 1-10.

[7] Bauder, R., da Rosa, R., & Khoshgoftaar, T. (2018). Identifying medicare provider fraud with unsupervised machine learning. In 2018 IEEE international conference on Information Reuse and Integration (IRI), IEEE, 18, 285-292.

[8] Bharne, S., & Bhaladhare, P. (2022). Investigating online dating fraud: An extensive review and analysis. In 2022 International Conference on Recent Trends in Microelectronics, Automation, Computing and Communications Systems (ICMACC), 141-147.

[9] Billah, M. M., Bhuiyan, M. N., & Akterujjaman, M. (2021). Unsupervised method of clustering and labeling of the online product based on reviews. International Journal of Modeling, Simulation, and Scientific Computing, 12(02), 2150017.

[10] Chiu, C. C., & Tsai, C. Y. (2004). A web services-based collaborative scheme for credit card fraud detection. In IEEE International Conference on e-Technology, e-Commerce and e-Service, 2004. EEE'04. 2004, 177-181.

[11] Cresci, S., Di Pietro, R., Petrocchi, M., Spognardi, A., & Tesconi, M. (2015). Fame for sale: Efficient detection of fake Twitter followers. Decision Support Systems, 80, 56-71.

[12] Mr DD Sarpate, "Design of Dual Band Microstrip for satellite Applications", 2nd International Conference on Recent Innovations in Engineering & Technology 2020

[13] Debener, J., Heinke, V., & Kriebel, J. (2023). Detecting insurance fraud using supervised and unsupervised machine learning. Journal of Risk and Insurance, 90(3), 743-768.

[14] Erşahin, B., Aktaş, Ö., Kılınç, D., & Akyol, C. (2017). Twitter fake account detection. In 2017 international conference on computer science and engineering (UBMK), 388-392.

[15] Usman, M., Zubair, M., Hussein, H. S., Wajid, M., Farrag, M., Ali, S. J., ... & Habeeb, M. S. (2021). Empirical mode decomposition for analysis and filtering of speech signals. IEEE Canadian Journal of Electrical and Computer Engineering, 44(3), 343-349.

[16] Singh, A., Gupta, M., Raj, A., Gupta, S. K., & Habeeb, M. S. (2020, December). TWDM-PON: The Enhanced PON for Triple Play Services. In 2020 5th IEEE International Conference on Recent Advances and Innovations in Engineering (ICRAIE) (pp. 1-5). IEEE.

[17] Naveen Sai Bommina, Nandipati Sai Akash, Uppu Lokesh, Dr. Hussain Syed, Dr. Syed Umar, "A Hybrid Optimization Framework for Enhancing IoT Security via AI-based Anomaly Detection", International Journal on Recent and Innovation Trends in Computing and Communication, (2023) ISSN: 2321-8169 Volume: 11 Issue: 3.

[18] Habeeb, M. S., & Babu, T. R. (2022). Network intrusion detection system: a survey on artificial intelligence‐based techniques. Expert Systems, 39(9), e13066

[19] Uppu Lokesh , Naveen Sai Bommina , Nandipati Sai Akash , Dr. Hussain Syed , Dr. Syed Umar. (2021). Deep Reinforcement Learning with Genetic Algorithm Tuning for Intrusion Detection in IoT Systems. International Journal of Communication Networks and Information Security (IJCNIS), 13(3), 582–595.

[20] Uppu Lokesh, Naveen Sai Bommina, Nandipati Sai Akash, Dr. Hussain Syed, Dr. Syed Umar, "Designing Energy-Efficient and Secure IoT Architectures Using Evolutionary Optimization Algorithms", International Journal of Applied Engineering & Technology, Vol. 4 No.2, September, 2022.

[21] Mr. Dikshendra Daulat Sarpate, and Dr. B.G Nagaraja, "CONVOLUTION NEURAL NETWORK-BASED SPEECH EMOTION RECOGNITION USING MFCCS", International Journal of Communication Networks and Information Security, 2023/12/10

[22] Naveen Sai Bommina , Nandipati Sai Akash, Uppu Lokesh , Dr. Hussain Syed , Dr. Syed Umar, "Multi-Objective Genetic Algorithms for Secure Routing and Data Privacy in IoT Networks", International Journal of Communication Networks and Information Security (IJCNIS), (2020), 12(3), 632–643.

[23] Nandipati Sai Akash, Naveen Sai Bommina, Uppu Lokesh, Hussain Syed, Syed Umar, "Optimized Block Chain-Enabled Security Mechanism for IoT Using Ant Colony Optimization", International Journal on Recent and Innovation Trends in Computing and Communication, (2023), 11(10), 1226–1233.

[24] Naveen Sai Bommina , Nandipati Sai Akash, Uppu Lokesh , Dr. Hussain Syed , Dr. Syed Umar, "Privacy-Preserving Federated Learning for IoT Devices with Secure Model Optimization", International Journal of Communication Networks and Information Security (IJCNIS), (2021), 13(2), 396–405.

[25] K Sankar, Divya Rohatgi, S Balakrishna Reddy, "COX Regressive Winsorized Correlated Convolutional Deep Belief Boltzmann Network for Covid-19 Prediction with Big Data", Grenze International Journal of Engineering & Technology (GIJET), Grenze ID: 01.GIJET.9.1.547, 2023.

[26] Naveen Sai Bommina, Uppu Lokesh, Nandipati Sai Akash, Dr. Hussain Syed, Dr. Syed Umar, "Optimized AI Models for Real -Time Cyberattack Detection in Smart Homes and Cities", International Journal of Applied Engineering & Technology, Vol. 4 No.1, June, 2022.

[27] K. Kartheeban, K. Kalyani, S. K. Bommavaram, D. Rohatgi, M. N. Kathiravan, and S. Saravanan, “Intelligent Deep Residual Network based Brain Tumor Detection and Classification,” in 2022 International Conference on Automation, Computing and Renewable Systems (ICACRS), Dec. 2022, pp. 785 –790.

[28] Umar, Syed, Bommina Naveen Sai, Nagineni Sai Lasya,Doppalapudi Asutosh, and LohithaRani. "Machine Learning based Sentiment Analysis of Product Reviews Using DeepEmbedding." Journal of Optoelectronics Laser 41, no. 6(2022): 108-113.

[29] Divya Rohatgi, Dr. Tulika Pandey, "Regression Test Selection Framework for Web Services", INTERNATIONAL JOURNAL OF SCIENTIFIC & TECHNOLOGY RESEARCH VOLUME 9, ISSUE 03, MARCH 2020.

How to cite this paper

R. Prasanth Reddy, Nagavelli Yogender Nath, Gattu Ramya, Syed Abdul Haq "Enhanced K-Means Clustering Implementation for Web-Based Suspicious Profiles Detection" Iconic Research And Engineering Journals Volume 7 Issue 8 2024 Page 548-555 https://doi.org/10.64388/IREV7I8-1712447
R. Prasanth Reddy, Nagavelli Yogender Nath, Gattu Ramya, Syed Abdul Haq "Enhanced K-Means Clustering Implementation for Web-Based Suspicious Profiles Detection" Iconic Research And Engineering Journals, vol. 7, no. 8, Feb. 2024, doi: https://doi.org/10.64388/IREV7I8-1712447
R. Prasanth Reddy, Nagavelli Yogender Nath, Gattu Ramya, Syed Abdul Haq (2024). Enhanced K-Means Clustering Implementation for Web-Based Suspicious Profiles Detection. Iconic Research And Engineering Journals, 7(8). doi: https://doi.org/10.64388/IREV7I8-1712447
R. Prasanth Reddy, Nagavelli Yogender Nath, Gattu Ramya, Syed Abdul Haq "Enhanced K-Means Clustering Implementation for Web-Based Suspicious Profiles Detection" Iconic Research And Engineering Journals, vol. 7, no. 8, Feb. 2024. Crossref, https://doi.org/10.64388/IREV7I8-1712447
@article{1712447,
      author = {R. Prasanth Reddy, Nagavelli Yogender Nath, Gattu Ramya, Syed Abdul Haq},
      title = {Enhanced K-Means Clustering Implementation for Web-Based Suspicious Profiles Detection},
      journal = {Iconic Research And Engineering Journals},
      year = {2024},
      volume = {7},
      number = {8},
      pages = {548-555},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1712447.pdf},
      abstract = {Fraud detection often requires a hybrid approach combining both human expertise and AI techniques. While AI models can process large amounts of data and detect intricate patterns, human intervention is sometimes necessary to refine rules, especially when dealing with ambiguous or doubtful cases. This is particularly important when distinguishing between fraudulent and legitimate behavior, as subtle differences can exist. Another challenge is the multilingual nature of the web. Fraud can occur across different languages and cultural contexts, making it difficult to build comprehensive fraud detection models that account for all variations. In cases where linguistic resources (such as parsers or language models) are not readily available, ML tools can be adapted to learn from the data itself. However, for such models to function effectively, it is crucial to develop a rich set of features, ensuring that the data is representative and inclusive of the varied aspects of web-based fraud. This paper compares the performance of various clustering algorithms and develops an enhanced k-means algorithm utilizing the Minkowski metric. This modified algorithm, applied to an unlabeled dataset from an online dating site, effectively clusters users into authentic and suspicious categories. Supervised machine learning techniques are subsequently employed to validate the proposed model using another labeled online dating fraud dataset from the US. This methodology underscores the potential of combining traditional clustering methods with machine learning techniques to enhance anomaly detection across various domains, offering a practical tool for administrators and security experts.},
      keywords = {Deep Learning, Approaches Web-based Fraud, Fraud Detection, k-means algorithm},
      month = {February},
      doi = {https://doi.org/10.64388/IREV7I8-1712447}
  }