International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1717944

1717944 Vol 9 · Issue 11 Download Paper

Email Header Analysis for Digital Forensics Using Machine Learning

Deepa B Dr. Balamurugan S

Subject area: Science,Engineering and Technology  ·  Area of research: Cybersecurity and Digital Forensics

DOI: https://doi.org/10.64388/IREV9I11-1717944

Abstract

Email communication remains the dominant vector for advanced cyber-attacks, including phishing, spoofing, and Business Email Compromise (BEC). Conventional email security mechanisms rely predominantly on content-based analysis encompassing Natural Language Processing (NLP), keyword filtering, and signature-based detection which are increasingly inadequate against modern adversarial techniques such as clean-text phishing, image-based payloads, and AI-generated deceptive messages. This paper presents a novel forensic-aware, machine learning-driven framework that shifts the analytical focus from email body content to Simple Mail Transfer Protocol (SMTP) header metadata. Email headers encode verifiable forensic information relay paths (Received fields), originating IP addresses, timestamp sequences, and authentication outcomes (SPF, DKIM, DMARC) that collectively represent a tamper-resistant record of an email's transmission behavior. Unlike body content, header metadata is structurally constrained and considerably more difficult for attackers to consistently manipulate across multiple relay nodes. The proposed framework introduces a comprehensive feature engineering pipeline that extracts temporal, network, topological, and authentication-level attributes from SMTP headers. These features are processed through an ensemble of machine learning models Random Forest (RF), Isolation Forest, and XGBoost enabling both supervised classification of known attack patterns and unsupervised detection of novel anomalies. A critical contribution of this research is the integration of forensic traceability with automated detection: the system reconstructs email transmission paths and preserves evidentiary artifacts suitable for digital forensic investigation and legal proceedings. Experimental evaluations on datasets derived from SpamAssassin and PhishTank repositories demonstrate that the proposed ensemble model achieves an F1-score of approximately 0.969 and an AUC-ROC of 0.983, representing a 5–8% improvement over content-only baseline models. False positive rates are simultaneously reduced from 11.2% to 3.1%, ensuring operational reliability in enterprise environments. This research establishes a scalable, intelligent, and forensically defensible paradigm for combating sophisticated email-based cyber-threats.

Keywords

Anomaly Detection, Digital Forensics, Email Spoofing, Ensemble Learning, Feature Engineering, Graph-Based Topology, SMTP Header Analysis, Digital Forensics, Phishing Detection

References

[1] E. Lochin, “STAMP: SMTP Server Topological Analysis by Message Headers Parsing,” in Proceedings of the IEEE Consumer Communications & Networking Conference (CCNC), Las Vegas, NV, USA, Jan. 2025.

[2] N. A. Mohammed, A. Hassan, and M. R. Yusof, “Recognizing Phishing in Emails Using NLP & ML Techniques,” in Proceedings of the IEEE International Conference on Computing Research (ICCR), 2025.

[3] M. Ajuluchukwu, O. Nwosu, and K. Adams, “Detecting Malicious Emails and Deceptive Websites Using Generative AI,” in Proceedings of the IEEE International Conference on Communication Networks (CICN), 2025.

[4] L. K. T. A. U. et al., “Email Armour: A Multi-Layered Email Defense Solution,” in Proceedings of the IEEE International Conference on Information Technology Research (ICITR), 2024.

[5] S. Abu-Nimeh, D. Nappa, X. Wang, and S. Nair, “A Comparison of Machine Learning Techniques for Phishing Detection,” in Proceedings of the Anti-Phishing Working Group eCrime Researchers Summit, Pittsburgh, PA, USA, 2007, pp. 60–69.

[6] M. Chandrasekaran, K. Narayanan, and S. Upadhyaya, “Phishing Email Detection Based on Structural Properties,” in Proceedings of the NYS Cyber Security Conference, Albany, NY, USA, 2006.

[7] J. Ma, L. Saul, S. Savage, and G. Voelker, “Learning to Detect Malicious URLs,” ACM Transactions on Intelligent Systems and Technology, vol. 2, no. 3, pp. 1–24, Apr. 2011.

[8] A. Blum, B. Wardman, T. Solorio, and G. Warner, “Lexical Feature Based Phishing URL Detection Using Online Learning,” in Proceedings of the ACM Workshop on Artificial Intelligence and Security (AISec), Chicago, IL, USA, 2010.

[9] F. Toolan and J. Carthy, “Feature Selection for Spam and Phishing Detection,” in Proceedings of the IEEE eCrime Researchers Summit, Dallas, TX, USA, 2010.

[10] R. Verma and K. Dyer, “On the Character of Phishing URLs: Accurate and Robust Statistical Learning Classifiers,” in Proceedings of the ACM Conference on Data and Application Security and Privacy (CODASPY), San Antonio, TX, USA, 2015.

[11] A. O. Adebowale, K. T. Lwin, and M. A. Hossain, “Intelligent Phishing Detection Scheme Using Deep Learning Algorithms,” Journal of Enterprise Information Management, vol. 35, no. 3, pp. 694–714, 2021.

[12] SpamAssassin Public Corpus. [Online]. Available: https://spamassassin.apache.org/publiccorpus/

[13] PhishTank. [Online]. Available: https://www.phishtank.com/

[14] TREC 2007 Spam Track. [Online]. Available: https://trec.nist.gov/data/spam.html

[15] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), San Francisco, CA, USA, 2016, pp. 785–794.

[16] F. T. Liu, K. M. Ting, and Z.-H. Zhou, “Isolation Forest,” in Proceedings of the IEEE International Conference on Data Mining (ICDM), Pisa, Italy, 2008, pp. 413–422.

[17] L. Breiman, “Random Forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.

[18] NIST, Guide to Integrating Forensic Techniques into Incident Response, NIST Special Publication 800-86, Gaithersburg, MD, USA, 2006.

How to cite this paper

Deepa B, Dr. Balamurugan S "Email Header Analysis for Digital Forensics Using Machine Learning" Iconic Research And Engineering Journals Volume 9 Issue 11 2026 Page 2589-2599 https://doi.org/10.64388/IREV9I11-1717944
Deepa B, Dr. Balamurugan S "Email Header Analysis for Digital Forensics Using Machine Learning" Iconic Research And Engineering Journals, vol. 9, no. 11, May. 2026, doi: https://doi.org/10.64388/IREV9I11-1717944
Deepa B, Dr. Balamurugan S (2026). Email Header Analysis for Digital Forensics Using Machine Learning. Iconic Research And Engineering Journals, 9(11). doi: https://doi.org/10.64388/IREV9I11-1717944
Deepa B, Dr. Balamurugan S "Email Header Analysis for Digital Forensics Using Machine Learning" Iconic Research And Engineering Journals, vol. 9, no. 11, May. 2026. Crossref, https://doi.org/10.64388/IREV9I11-1717944
@article{1717944,
      author = {Deepa B, Dr. Balamurugan S},
      title = {Email Header Analysis for Digital Forensics Using Machine Learning},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {11},
      pages = {2589-2599},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1717944.pdf},
      abstract = {Email communication remains the dominant vector for advanced cyber-attacks, including phishing, spoofing, and Business Email Compromise (BEC). Conventional email security mechanisms rely predominantly on content-based analysis encompassing Natural Language Processing (NLP), keyword filtering, and signature-based detection which are increasingly inadequate against modern adversarial techniques such as clean-text phishing, image-based payloads, and AI-generated deceptive messages. This paper presents a novel forensic-aware, machine learning-driven framework that shifts the analytical focus from email body content to Simple Mail Transfer Protocol (SMTP) header metadata. Email headers encode verifiable forensic information relay paths (Received fields), originating IP addresses, timestamp sequences, and authentication outcomes (SPF, DKIM, DMARC) that collectively represent a tamper-resistant record of an email's transmission behavior. Unlike body content, header metadata is structurally constrained and considerably more difficult for attackers to consistently manipulate across multiple relay nodes. The proposed framework introduces a comprehensive feature engineering pipeline that extracts temporal, network, topological, and authentication-level attributes from SMTP headers. These features are processed through an ensemble of machine learning models Random Forest (RF), Isolation Forest, and XGBoost enabling both supervised classification of known attack patterns and unsupervised detection of novel anomalies. A critical contribution of this research is the integration of forensic traceability with automated detection: the system reconstructs email transmission paths and preserves evidentiary artifacts suitable for digital forensic investigation and legal proceedings. Experimental evaluations on datasets derived from SpamAssassin and PhishTank repositories demonstrate that the proposed ensemble model achieves an F1-score of approximately 0.969 and an AUC-ROC of 0.983, representing a 5–8% improvement over content-only baseline models. False positive rates are simultaneously reduced from 11.2% to 3.1%, ensuring operational reliability in enterprise environments. This research establishes a scalable, intelligent, and forensically defensible paradigm for combating sophisticated email-based cyber-threats.},
      keywords = {Anomaly Detection, Digital Forensics, Email Spoofing, Ensemble Learning, Feature Engineering, Graph-Based Topology, SMTP Header Analysis, Digital Forensics, Phishing Detection},
      month = {May},
      doi = {https://doi.org/10.64388/IREV9I11-1717944}
  }