Home / Current Issue / Paper 1710433
Comparative Analysis of Machine Learning and Deep Learning Approaches for Phishing Email Detection Using Natural Language Processing
Subject area: Science,Engineering and Technology · Area of research: Cybersecurity
Abstract
One of the biggest threats to cybersecurity today is phishing emails that find and take advantage of human weaknesses, and they fool common detection methods. This paper reserves a contrast between Deep Learning (DL) and Machine Learning (ML) of identifying phishing email through Natural Language Processing (NLP). A publicly available dataset in Kaggle (82,797 emails of which 42,890 emails are phishing and 39,595 emails are non-phishing) was examined. To standardize the texts, preprocessing methods such as tokenization, stop-word elimination, and lemmatization were undertaken before the either Term Frequency ? Inverse Document Frequency (TF-IDF) to create features in the ML models and bidirectional Encoder Representations from Transformers ? Long Short-Term Memory model (BERT-LSTM). ML models that were used were Random Forest (RF) and Support Vector Machine (SVM) and the DL model was a BERT one. Analysis showed there were different trade-offs of the approaches. TF-IDF and ML performed well and have lesser CPU load, which can be used in a situation where resources are scarce. Specifically, the Random Forest performed well considering its power of ensemble, SVM with the linearity kernel in dealing with high dimensions. On the other hand, BERT-LSTM model has proven to be more accurate because it embraces contextual and semantics of email text, at the expense of greater computing burden. The results support the argument that the selection of the technique must favor accuracy and availability of resources. Though ML based on TF-IDF will have a lightweight and practical solution, DL based on BERT- LSTM presents sophisticated context-related insight to phishing detection solutions when it comes to applications that involve high stakes.
Keywords
Cybersecurity, Deep learning, Machine learning, Natural language processing, Phishing detection.
References
[1] Adwan, Y., & Abuhasan, A. (2016). An intelligent classification model for phishing email detection. arXiv.
[2] Al-Falahi, A. S., Al-Omaishi, N., & Al-Zubaidi, A. A. (2021). A review of deep learning techniques for phishing email detection. Journal of Cyber Security and Mobility, 10(2), 291–320.
[3] Amazon Web Services. (n.d.). What is natural language processing? Amazon. Retrieved August 18, 2025, from https://aws.amazon.com/what-is/nlp/
[4] ESET Editorial Team. (2023). Vishing, smishing, and phishing: How to arm yourself against social engineering attacks. ESET. Retrieved August 19, 2025, from https://www.eset.com/blog/en/business-topics/threat-landscape/social-engineering-vishing-phishing/
[5] Gupta, M., Sharma, S., & Agrawal, A. (2017). Phishing attack detection using machine learning techniques: A review. International Journal of Computer Applications, 163(8), 1–6.
[6] IBM. (2024). What is natural language processing? IBM. Retrieved August 18, 2025, from https://www.ibm.com/think/topics/natural-language-processing
[7] Imperva. (2023). Phishing attack: Scam definition, types, and examples. Imperva. Retrieved August 18, 2025, from https://www.imperva.com/learn/application-security/phishing-attack-scam/
[8] Jakobsson, M., & Myers, S. (2006). Phishing and countermeasures: Understanding the increasing threat of online identity theft. Wiley.
[9] Khadka, K., Ullah, A. B., Ma, W., & Martinez Marroquin, E. (2024). A survey on the principles of persuasion as a social engineering strategy in phishing. arXiv. Retrieved August 19, 2025.
[10] Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
[11] Mittal, P., Singh, H., & Sood, S. K. (2021). Phishing detection using machine learning and deep learning with TF-IDF. International Journal of Information and Computer Security, 15(1), 1–15.
[12] O’Gorman, E. (2007). Security and privacy in the age of pervasive computing. Artech House.
[13] Samarthrao, S., & Rohokale, V. (2022). Phishing detection leveraging machine learning and deep learning: A review. Electronics, 12(21), 4545.
[14] Wikipedia. (2025). Natural language processing. In Wikipedia. Retrieved August 18, 2025, from https://en.wikipedia.org/wiki/Natural_language_processing
[15] Wikipedia. (2025). Phishing. In Wikipedia. Retrieved August 18, 2025, from https://en.wikipedia.org/wiki/Phishing
[16] Zabihimayvan, M., & Daremey, F. (2020). Phishing email detection using machine learning algorithms and a hybrid approach. Journal of Computer Science and Technology, 35(6), 1339–1355.
[17] Zahid, M., Ramanna, V., Kenchamma, R. H., & Basapur, S. B. (2021). Applying machine learning and natural language processing to detect phishing email. Computers & Security, 110, 102414.
How to cite this paper
@article{1710433,
author = {BUOYE, Peter. Adewuyi, AKINBOLA, Sherifat. Morenike},
title = {Comparative Analysis of Machine Learning and Deep Learning Approaches for Phishing Email Detection Using Natural Language Processing},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {3},
pages = {165-170},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1710433.pdf},
abstract = {One of the biggest threats to cybersecurity today is phishing emails that find and take advantage of human weaknesses, and they fool common detection methods. This paper reserves a contrast between Deep Learning (DL) and Machine Learning (ML) of identifying phishing email through Natural Language Processing (NLP). A publicly available dataset in Kaggle (82,797 emails of which 42,890 emails are phishing and 39,595 emails are non-phishing) was examined. To standardize the texts, preprocessing methods such as tokenization, stop-word elimination, and lemmatization were undertaken before the either Term Frequency ? Inverse Document Frequency (TF-IDF) to create features in the ML models and bidirectional Encoder Representations from Transformers ? Long Short-Term Memory model (BERT-LSTM). ML models that were used were Random Forest (RF) and Support Vector Machine (SVM) and the DL model was a BERT one. Analysis showed there were different trade-offs of the approaches. TF-IDF and ML performed well and have lesser CPU load, which can be used in a situation where resources are scarce. Specifically, the Random Forest performed well considering its power of ensemble, SVM with the linearity kernel in dealing with high dimensions. On the other hand, BERT-LSTM model has proven to be more accurate because it embraces contextual and semantics of email text, at the expense of greater computing burden. The results support the argument that the selection of the technique must favor accuracy and availability of resources. Though ML based on TF-IDF will have a lightweight and practical solution, DL based on BERT- LSTM presents sophisticated context-related insight to phishing detection solutions when it comes to applications that involve high stakes.},
keywords = {Cybersecurity, Deep learning, Machine learning, Natural language processing, Phishing detection.},
month = {September},
}