International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1723012

1723012 Vol 10 · Issue 3 Download Paper

Machine Learning-Based Phishing Detection Techniques: A Systematic Literature Review

Ekpo Precious Emmanuel Solomon Ahiba Assoc. Prof. Vivian Nwaocha Eru Akwuma Nathaniel Ojima Gabriella Ob'lama

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

Abstract

Phishing remains a persistent cybersecurity threat because attackers combine social engineering with rapidly changing emails, websites, URLs, text messages, QR codes, and mobile channels. Static blacklists and manually defined rules are increasingly inadequate for previously unseen campaigns. This study systematically reviews machine learning-based phishing detection research published between 2020 and 2026. A PRISMA-guided protocol was used to identify, screen, appraise, and synthesise 50 studies. The review examined the techniques employed, their reported effectiveness, the datasets and evaluation metrics used, and the principal limitations and emerging research directions. Five technique families were identified: traditional machine learning, ensemble learning, deep learning, transformer and large-language-model approaches, and explainable or hybrid intelligent systems. Traditional classifiers remain attractive for efficiency and interpretability, while ensemble, deep, transformer, multimodal, and continual-learning methods increasingly address complex features, contextual language, concept drift, and cross-channel attacks. PhishTank, OpenPhish, the UCI Phishing Websites dataset, ISCX-URL2016, SpamAssassin, Enron, and related repositories were commonly used. Accuracy was reported in 48 of the 50 studies, followed by precision (43), recall (42), and F1-score (41), whereas ROC-AUC, Matthews correlation coefficient, and error-rate measures were less frequent. Dataset ageing, class imbalance, weak cross-dataset generalisation, adversarial manipulation, computational cost, and limited explainability remain major barriers. Practical progress depends on adaptive, lightweight, explainable, and robust models evaluated with current datasets and standardised multi-metric protocols.

Keywords

Cybersecurity, deep learning, explainable artificial intelligence, large language models, machine learning, phishing detection, systematic literature review, transformer models

How to cite this paper

Ekpo Precious Emmanuel, Solomon Ahiba, Assoc. Prof. Vivian Nwaocha, Eru Akwuma Nathaniel, Ojima Gabriella Ob'lama "Machine Learning-Based Phishing Detection Techniques: A Systematic Literature Review" Iconic Research And Engineering Journals Volume 10 Issue 3 2026 Page 1196-1206
Ekpo Precious Emmanuel, Solomon Ahiba, Assoc. Prof. Vivian Nwaocha, Eru Akwuma Nathaniel, Ojima Gabriella Ob'lama "Machine Learning-Based Phishing Detection Techniques: A Systematic Literature Review" Iconic Research And Engineering Journals, vol. 10, no. 3, Sep. 2026
Ekpo Precious Emmanuel, Solomon Ahiba, Assoc. Prof. Vivian Nwaocha, Eru Akwuma Nathaniel, Ojima Gabriella Ob'lama (2026). Machine Learning-Based Phishing Detection Techniques: A Systematic Literature Review. Iconic Research And Engineering Journals, 10(3).
Ekpo Precious Emmanuel, Solomon Ahiba, Assoc. Prof. Vivian Nwaocha, Eru Akwuma Nathaniel, Ojima Gabriella Ob'lama "Machine Learning-Based Phishing Detection Techniques: A Systematic Literature Review" Iconic Research And Engineering Journals, vol. 10, no. 3, Sep. 2026.
@article{1723012,
      author = {Ekpo Precious Emmanuel, Solomon Ahiba, Assoc. Prof. Vivian Nwaocha, Eru Akwuma Nathaniel, Ojima Gabriella Ob'lama},
      title = {Machine Learning-Based Phishing Detection Techniques: A Systematic Literature Review},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {10},
      number = {3},
      pages = {1196-1206},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1723012.pdf},
      abstract = {Phishing remains a persistent cybersecurity threat because attackers combine social engineering with rapidly changing emails, websites, URLs, text messages, QR codes, and mobile channels. Static blacklists and manually defined rules are increasingly inadequate for previously unseen campaigns. This study systematically reviews machine learning-based phishing detection research published between 2020 and 2026. A PRISMA-guided protocol was used to identify, screen, appraise, and synthesise 50 studies. The review examined the techniques employed, their reported effectiveness, the datasets and evaluation metrics used, and the principal limitations and emerging research directions. Five technique families were identified: traditional machine learning, ensemble learning, deep learning, transformer and large-language-model approaches, and explainable or hybrid intelligent systems. Traditional classifiers remain attractive for efficiency and interpretability, while ensemble, deep, transformer, multimodal, and continual-learning methods increasingly address complex features, contextual language, concept drift, and cross-channel attacks. PhishTank, OpenPhish, the UCI Phishing Websites dataset, ISCX-URL2016, SpamAssassin, Enron, and related repositories were commonly used. Accuracy was reported in 48 of the 50 studies, followed by precision (43), recall (42), and F1-score (41), whereas ROC-AUC, Matthews correlation coefficient, and error-rate measures were less frequent. Dataset ageing, class imbalance, weak cross-dataset generalisation, adversarial manipulation, computational cost, and limited explainability remain major barriers. Practical progress depends on adaptive, lightweight, explainable, and robust models evaluated with current datasets and standardised multi-metric protocols.},
      keywords = {Cybersecurity, deep learning, explainable artificial intelligence, large language models, machine learning, phishing detection, systematic literature review, transformer models},
      month = {September},
  }