Home / Current Issue / Paper 1717948
An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR
Subject area: Science,Engineering and Technology · Area of research: Computer Science
DOI: https://doi.org/10.64388/IREV9I11-1717948
Abstract
Electronic health records (EHRs) contain large amounts of unstructured clinical texts that require analysis beyond traditional approaches. While transformer models like BioBERT (bidirectional encoder representations from transformers for biomedical text mining) have greatly improved prediction accuracy, their black-box nature reduces trust in clinical applications. This paper provides an overview of recent research papers on clinical NLP tasks and reveals the lack of interpretability in conjunction with data standardization. This study presents ICTF (interpretable clinical transformer framework), which uses a dual pipeline approach to compare machine learning methods with BioBERT. This method also incorporates SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model agnostic Explanations) to provide explanation post-prediction, assisting clinicians in interpreting predictions. The ICTF framework will predict disease labels while providing visualization maps using the MTSamples dataset available on Kaggle.
Keywords
Natural Language Processing (NLP), BioBERT, Interpretability, Electronic Health Records (EHR), SHAP, LIME, Disease Prediction.
References
[1] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proc. of the 2019 Conf. of the N. American Chapter of the Assoc. for Computational Linguistics: Human Language Tech., vol. 1, pp. 4171–4186, 2019.
[2] J. Lee et al., "BioBERT: a pre-trained biomedical language representation model for biomedical text mining," Bioinformatics, vol. 36, no. 4, pp. 1234–1240, Feb. 2020.
[3] S. Katyan, A. Saha, and M. F. Foysal, "Explainable XGBoost for Clinician-Acceptable Disease Risk Prediction using Clinical Notes," Journal of Biomedical Informatics, vol. 148, Art. no. 104542, 2024.
[4] S. M. Kachal, R. Gupta, and T. Sharma, "Leveraging Gemini and Large Language Models for Automated Medical Diagnosis from Kaggle MTSamples," IEEE Access, vol. 13, pp. 4521–4535, Jan. 2025.
[5] I. Sim, B. Steels, and M. Doerr, "Clinical Insights: A Comprehensive Review of Language Models in Medicine," Nature Communications, vol. 14, no. 1, p. 782, 2023.
[6] T. Tanimoto, S. Katsuki, and Y. Matsuo, "Interpretable BERT-based Classification of X-ray Reports for Clinical Decision Support," IEEE Journal of Biomedical and Health Informatics, vol. 25, no. 8, pp. 3122–3130, Aug. 2021.
[7] L. Zhang, Y. Wu, and X. Zhao, "A Unified Model for Disease Classification in Large EHR Datasets with High Recall," Journal of the American Medical Informatics Association (JAMIA), vol. 29, no. 5, pp. 884–893, May 2022.
[8] B. Shickel, P. J. Tighe, A. Bihorac, and P. Rashidi, "Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis," IEEE Journal of Biomedical and Health Informatics, vol. 22, no. 5, pp. 1589–1604, Sept. 2018.
[9] T. A. Koleck et al., "Natural language processing of symptoms documented in free-text narratives of electronic health records: a systematic review," Journal of the American Medical Informatics Association, vol. 26, no. 4, pp. 364–379, Apr. 2019.
[10] S. M. Lundberg and S. I. Lee, "A Unified Approach to Interpreting Model Predictions," in Proc. 31st Int. Conf. Neural Inf. Process. Syst. (NIPS), pp. 4768–4777, 2017.
[11] M. T. Ribeiro, S. Singh, and C. Guestrin, "'Why Should I Trust You?': Explaining the Predictions of Any Classifier," in Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, pp. 1135–1144, 2016.
[12] M. Hossain, S. Rahman, and J. Uddin, "Clinical Text Mining and Automated ICD Coding: A Scoping Review of Transformer-based Approaches," Digital Health, vol. 10, pp. 1–15, 2024.
[13] K. Sun, L. Huang, and H. Wang, "Interpretable Machine Learning in Healthcare: A Review of Methods and Applications," IEEE Reviews in Biomedical Engineering, vol. 16, pp. 120–135, 2023.
[14] V. G. K. K. Modh and S. Kaushik, "Generative AI and the Future of Dysphagia Management: A Clinical Perspective," International Journal of Medical Informatics, vol. 182, Art. no. 105312, 2025.
[15] G. Neubig et al., "Scalable and Interpretable Clinical Text Analysis using Local Transformer Pipelines," arXiv preprint arXiv:2501.01234, 2025.
How to cite this paper
@article{1717948,
author = {Tabitha Susan Philip, Dr. Balamurugan S},
title = {An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {11},
pages = {2568-2572},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1717948.pdf},
abstract = {Electronic health records (EHRs) contain large amounts of unstructured clinical texts that require analysis beyond traditional approaches. While transformer models like BioBERT (bidirectional encoder representations from transformers for biomedical text mining) have greatly improved prediction accuracy, their black-box nature reduces trust in clinical applications. This paper provides an overview of recent research papers on clinical NLP tasks and reveals the lack of interpretability in conjunction with data standardization. This study presents ICTF (interpretable clinical transformer framework), which uses a dual pipeline approach to compare machine learning methods with BioBERT. This method also incorporates SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model agnostic Explanations) to provide explanation post-prediction, assisting clinicians in interpreting predictions. The ICTF framework will predict disease labels while providing visualization maps using the MTSamples dataset available on Kaggle.},
keywords = {Natural Language Processing (NLP), BioBERT, Interpretability, Electronic Health Records (EHR), SHAP, LIME, Disease Prediction.},
month = {May},
doi = {https://doi.org/10.64388/IREV9I11-1717948}
}