International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1717948

1717948 Vol 9 · Issue 11 Download Paper

An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR

Tabitha Susan Philip Dr. Balamurugan S

Subject area: Science,Engineering and Technology  ·  Area of research: Computer Science

DOI: https://doi.org/10.64388/IREV9I11-1717948

Abstract

Electronic health records (EHRs) contain large amounts of unstructured clinical texts that require analysis beyond traditional approaches. While transformer models like BioBERT (bidirectional encoder representations from transformers for biomedical text mining) have greatly improved prediction accuracy, their black-box nature reduces trust in clinical applications. This paper provides an overview of recent research papers on clinical NLP tasks and reveals the lack of interpretability in conjunction with data standardization. This study presents ICTF (interpretable clinical transformer framework), which uses a dual pipeline approach to compare machine learning methods with BioBERT. This method also incorporates SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model agnostic Explanations) to provide explanation post-prediction, assisting clinicians in interpreting predictions. The ICTF framework will predict disease labels while providing visualization maps using the MTSamples dataset available on Kaggle.

Keywords

Natural Language Processing (NLP), BioBERT, Interpretability, Electronic Health Records (EHR), SHAP, LIME, Disease Prediction.

References

[1] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proc. of the 2019 Conf. of the N. American Chapter of the Assoc. for Computational Linguistics: Human Language Tech., vol. 1, pp. 4171–4186, 2019.

[2] J. Lee et al., "BioBERT: a pre-trained biomedical language representation model for biomedical text mining," Bioinformatics, vol. 36, no. 4, pp. 1234–1240, Feb. 2020.

[3] S. Katyan, A. Saha, and M. F. Foysal, "Explainable XGBoost for Clinician-Acceptable Disease Risk Prediction using Clinical Notes," Journal of Biomedical Informatics, vol. 148, Art. no. 104542, 2024.

[4] S. M. Kachal, R. Gupta, and T. Sharma, "Leveraging Gemini and Large Language Models for Automated Medical Diagnosis from Kaggle MTSamples," IEEE Access, vol. 13, pp. 4521–4535, Jan. 2025.

[5] I. Sim, B. Steels, and M. Doerr, "Clinical Insights: A Comprehensive Review of Language Models in Medicine," Nature Communications, vol. 14, no. 1, p. 782, 2023.

[6] T. Tanimoto, S. Katsuki, and Y. Matsuo, "Interpretable BERT-based Classification of X-ray Reports for Clinical Decision Support," IEEE Journal of Biomedical and Health Informatics, vol. 25, no. 8, pp. 3122–3130, Aug. 2021.

[7] L. Zhang, Y. Wu, and X. Zhao, "A Unified Model for Disease Classification in Large EHR Datasets with High Recall," Journal of the American Medical Informatics Association (JAMIA), vol. 29, no. 5, pp. 884–893, May 2022.

[8] B. Shickel, P. J. Tighe, A. Bihorac, and P. Rashidi, "Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis," IEEE Journal of Biomedical and Health Informatics, vol. 22, no. 5, pp. 1589–1604, Sept. 2018.

[9] T. A. Koleck et al., "Natural language processing of symptoms documented in free-text narratives of electronic health records: a systematic review," Journal of the American Medical Informatics Association, vol. 26, no. 4, pp. 364–379, Apr. 2019.

[10] S. M. Lundberg and S. I. Lee, "A Unified Approach to Interpreting Model Predictions," in Proc. 31st Int. Conf. Neural Inf. Process. Syst. (NIPS), pp. 4768–4777, 2017.

[11] M. T. Ribeiro, S. Singh, and C. Guestrin, "'Why Should I Trust You?': Explaining the Predictions of Any Classifier," in Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, pp. 1135–1144, 2016.

[12] M. Hossain, S. Rahman, and J. Uddin, "Clinical Text Mining and Automated ICD Coding: A Scoping Review of Transformer-based Approaches," Digital Health, vol. 10, pp. 1–15, 2024.

[13] K. Sun, L. Huang, and H. Wang, "Interpretable Machine Learning in Healthcare: A Review of Methods and Applications," IEEE Reviews in Biomedical Engineering, vol. 16, pp. 120–135, 2023.

[14] V. G. K. K. Modh and S. Kaushik, "Generative AI and the Future of Dysphagia Management: A Clinical Perspective," International Journal of Medical Informatics, vol. 182, Art. no. 105312, 2025.

[15] G. Neubig et al., "Scalable and Interpretable Clinical Text Analysis using Local Transformer Pipelines," arXiv preprint arXiv:2501.01234, 2025.

How to cite this paper

Tabitha Susan Philip, Dr. Balamurugan S "An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR" Iconic Research And Engineering Journals Volume 9 Issue 11 2026 Page 2568-2572 https://doi.org/10.64388/IREV9I11-1717948
Tabitha Susan Philip, Dr. Balamurugan S "An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR" Iconic Research And Engineering Journals, vol. 9, no. 11, May. 2026, doi: https://doi.org/10.64388/IREV9I11-1717948
Tabitha Susan Philip, Dr. Balamurugan S (2026). An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR. Iconic Research And Engineering Journals, 9(11). doi: https://doi.org/10.64388/IREV9I11-1717948
Tabitha Susan Philip, Dr. Balamurugan S "An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR" Iconic Research And Engineering Journals, vol. 9, no. 11, May. 2026. Crossref, https://doi.org/10.64388/IREV9I11-1717948
@article{1717948,
      author = {Tabitha Susan Philip, Dr. Balamurugan S},
      title = {An Interpretable Clinical Transformer Framework (ICTF) for Disease Prediction using EHR},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {11},
      pages = {2568-2572},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1717948.pdf},
      abstract = {Electronic health records (EHRs) contain large amounts of unstructured clinical texts that require analysis beyond traditional approaches. While transformer models like BioBERT (bidirectional encoder representations from transformers for biomedical text mining) have greatly improved prediction accuracy, their black-box nature reduces trust in clinical applications.  This paper provides an overview of recent research papers on clinical NLP tasks and reveals the lack of interpretability in conjunction with data standardization. This study presents ICTF (interpretable clinical transformer framework), which uses a dual pipeline approach to compare machine learning methods with BioBERT. This method also incorporates SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model agnostic Explanations) to provide explanation post-prediction, assisting clinicians in interpreting predictions. The ICTF framework will predict disease labels while providing visualization maps using the MTSamples dataset available on Kaggle.},
      keywords = {Natural Language Processing (NLP), BioBERT, Interpretability, Electronic Health Records (EHR), SHAP, LIME, Disease Prediction.},
      month = {May},
      doi = {https://doi.org/10.64388/IREV9I11-1717948}
  }