International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1723866

1723866 Vol 10 · Issue 4 Download Paper

Comparative Analysis of Different Machine Learning Models to Recognise Spam and Benign Mail Using URL

Debdutta Banerjee Soumendu Banerjee

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

Abstract

Phishing attacks are the most prevalent and concerning cybersecurity threats. Phishing typically involves exploiting malicious URLs to deceive users into revealing sensitive information or accessing fraudulent websites. The increasing rates of phishing techniques require the development of accurate and automated detection mechanisms. This study proposes a machine learning–based phishing URL detection framework that utilises Term Frequency–Inverse Document Frequency (TF-IDF) used here for feature extraction from URL strings. A comprehensive comparative analysis was conducted using four widely adopted supervised learning algorithms: Logistic Regression, Linear Support Vector Machine (SVM), Random Forest, and Extreme Gradient Boosting (XGBoost). Experiments were performed on a large-scale phishing URL dataset comprising 871,590 labelled URL samples. The models were evaluated using standard performance metrics, including accuracy, precision, recall, and F1-score. The experimental results indicate that XGBoost, when combined with TF-IDF features, consistently outperformed the other classifiers, achieving the highest detection accuracy and superior classification performance. The findings demonstrate the effectiveness of integrating TF-IDF-based lexical feature extraction with ensemble learning techniques for phishing URL classification. The proposed framework offers a scalable and reliable solution for real-time phishing detection and can be integrated into web browsers, email filtering systems, and cybersecurity applications to enhance protection against evolving phishing attacks.

Keywords

phishing URL detection; machine learning; TF-IDF; XGBoost; cybersecurity; malicious URL classification

References

[1] APWG. Phishing activity trends report. Anti-Phishing Working Group; 2024.

[2] Verma R, Dyer K. On the character of phishing URLs: Accurate and robust statistical learning classifiers. In: Proceedings of the ACM CODASPY. 2015.

[3] Khonji M, Iraqi Y, Jones A. Phishing detection: A literature survey. IEEE Communications Surveys & Tutorials. 2013;15(4):2091–2121.

[4] Marchal S, François J, State R, Engel T. PhishStorm: Detecting phishing with streaming analytics. IEEE Transactions on Network and Service Management. 2014;11(4):458–471.

[5] Sahingoz Y, Buber E, Demir O, Diri B. Machine learning based phishing detection from URLs. Expert Systems with Applications. 2019;117:345–357.

[6] James C, Joseph A, Balakrishnan V. Phishing detection using machine learning techniques: A review. Journal of Information Security and Applications. 2021.

[7] Chen T, Guestrin C. XGBoost: A scalable tree boosting system. In: Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. p. 785–794.

[8] Salton G, Buckley C. Term-weighting approaches in automatic text retrieval. Information Processing & Management. 1988;24(5):513–523.

[9] Sanchez-Paniagua M, Fidalgo E, Alegre E, Al-Nabki W, Gonzalez-Castro V. Phishing URL detection: A real-case scenario through login URLs. IEEE Access. 2022;10:42949–42960.

[10] Le H, Pham Q, Sahoo D, Hoi SCH. URLNet: Learning a URL representation with deep learning for malicious URL detection [preprint]. arXiv. 2018. arXiv:1802.03162

[11] Kustiawan YA, Ghauth KI. Evaluating the impact of feature engineering in phishing URL detection: A comparative study of URL, HTML, and derived features. IEEE Access. 2025;13:126756–126768.

[12] Chen T, Guestrin C. XGBoost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). 2016. p. 785–794.

[13] Asiri S, Xiao Y, Alzahrani S, Li S, Li T. A survey of intelligent detection designs of HTML URL phishing attacks. IEEE Access. 2023;11:6421–6443.

[14] Chanda T, Mondal B, Das S, Banerjee S, Mondal B. An explainable AI-based hybrid SVM–CNN ensemble model implementation for intelligent email spam detection. In: 2026 International Conference on Emerging Systems and Intelligent Computing (ESIC); Bhubaneswar, India. 2026. p. 240–245. https://doi.org/10.1109/ESIC68176.2026.11496294

[15] Chanda T, Mondal B, Das S, Banerjee S, Mondal B. Temporal analysis based exploratory data analysis for phishing email detection via machine learning. In: Chatterjee B, Kothapalli K, Mittal N, Natarajan AM, Singh D, editors. Distributed computing and intelligent technology (ICDCIT 2026). Cham: Springer; 2026. (Lecture Notes in Computer Science; vol. 16420). https://doi.org/10.1007/978-3-032-16632-6_21

How to cite this paper

Debdutta Banerjee, Soumendu Banerjee "Comparative Analysis of Different Machine Learning Models to Recognise Spam and Benign Mail Using URL" Iconic Research And Engineering Journals Volume 10 Issue 4 2026 Page 1273-1283
Debdutta Banerjee, Soumendu Banerjee "Comparative Analysis of Different Machine Learning Models to Recognise Spam and Benign Mail Using URL" Iconic Research And Engineering Journals, vol. 10, no. 4, Oct. 2026
Debdutta Banerjee, Soumendu Banerjee (2026). Comparative Analysis of Different Machine Learning Models to Recognise Spam and Benign Mail Using URL. Iconic Research And Engineering Journals, 10(4).
Debdutta Banerjee, Soumendu Banerjee "Comparative Analysis of Different Machine Learning Models to Recognise Spam and Benign Mail Using URL" Iconic Research And Engineering Journals, vol. 10, no. 4, Oct. 2026.
@article{1723866,
      author = {Debdutta Banerjee, Soumendu Banerjee},
      title = {Comparative Analysis of Different Machine Learning Models to Recognise Spam and Benign Mail Using URL},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {10},
      number = {4},
      pages = {1273-1283},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1723866.pdf},
      abstract = {Phishing attacks are the most prevalent and concerning cybersecurity threats. Phishing typically involves exploiting malicious URLs to deceive users into revealing sensitive information or accessing fraudulent websites. The increasing rates of phishing techniques require the development of accurate and automated detection mechanisms. This study proposes a machine learning–based phishing URL detection framework that utilises Term Frequency–Inverse Document Frequency (TF-IDF) used here for feature extraction from URL strings. A comprehensive comparative analysis was conducted using four widely adopted supervised learning algorithms: Logistic Regression, Linear Support Vector Machine (SVM), Random Forest, and Extreme Gradient Boosting (XGBoost). Experiments were performed on a large-scale phishing URL dataset comprising 871,590 labelled URL samples. The models were evaluated using standard performance metrics, including accuracy, precision, recall, and F1-score. The experimental results indicate that XGBoost, when combined with TF-IDF features, consistently outperformed the other classifiers, achieving the highest detection accuracy and superior classification performance. The findings demonstrate the effectiveness of integrating TF-IDF-based lexical feature extraction with ensemble learning techniques for phishing URL classification. The proposed framework offers a scalable and reliable solution for real-time phishing detection and can be integrated into web browsers, email filtering systems, and cybersecurity applications to enhance protection against evolving phishing attacks.},
      keywords = {phishing URL detection; machine learning; TF-IDF; XGBoost; cybersecurity; malicious URL classification},
      month = {October},
  }