Home / Current Issue / Paper 1711765
A Unified Multi-Modal Transformer Framework for Synergistic Cancer Diagnosis
Subject area: Science,Engineering and Technology · Area of research: Multi-modal Learning, Precision Medicine
Abstract
Early cancer diagnosis is critical for improving patient outcomes but is challenged by the disease?s profound heterogeneity. This paper introduces a unified, AI-powered framework that synergistically integrates histopathology, genomics, and proteomics data to enhance early cancer detection. Our architecture features a novel multi-transformer model with dedicated Vision and Genomic Transformers to encode modality-specific features, which are then fused by a cross-modal attention transformer. This intermediate fusion strategy enables the model to learn intricate genotype-phenotype correlations often missed by traditional methods. Validated on cohorts from The Cancer Genome Atlas (TCGA), our framework demonstrates a significant improvement in diagnostic performance over single-modality baselines. We also incorporate Explainable AI (XAI) techniques to ensure model transparency, a crucial step for clinical adoption. The framework serves as both a powerful diagnostic tool and a hypothesis- generation engine, uncovering novel biomarkers from complex multi-modal data and advancing computational pathology and personalized medicine.
Keywords
Multi-Modal Learning, Transformers, Histopathology, Genomics, Proteomics, Explainable AI, Early Cancer Diagnosis
References
[1] K. D. McCombe, S. G. Craig, A. V. Pulsawatdi, et al., “HistoClean: Open-source software for histological image pre-processing and aug- mentation to improve development of robust convolutional neural net- works,” Computational and Structural Biotechnology Journal, vol. 19, pp. 4840–4853, 2021.
[2] S. S. Band, A. Yarahmadi, C.-C. Hsu, et al., “Application of explain- able artificial intelligence in medical health: A systematic review of interpretability methods,” Informatics in Medicine Unlocked, vol. 40, p. 101286, 2023.
[3] X. Li, M. Li, P. Yan, et al., “Deep Learning Attention Mechanism in Medical Image Analysis: Basics and Beyonds,” International Journal of Network Dynamics and Intelligence, vol. 2, no. 1, pp. 93–116, 2023.
[4] Z. Zhou, X. Feng, L. Huang, et al., “From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems,” arXiv preprint arXiv:2503.01424, 2025.
[5] D. Wilimitis and C. G. Walsh, “Practical Considerations and Applied Examples of Cross-Validation for Model Development and Evaluation in Health Care: Tutorial,” JMIR AI, vol. 2, p. e49023, 2023.
[6] J.-K. He´riche´, S. Alexander, and J. Ellenberg, “Integrating Imaging and Omics: Computational Methods and Challenges,” Annual Review of Biomedical Data Science, vol. 2, pp. 175–197, 2019.
[7] C.-H. Liu, C.-F. Tsai, K.-L. Sue, and M.-W. Huang, “The Feature Selection Effect on Missing Value Imputation of Medical Datasets,” Applied Sciences, vol. 10, no. 7, p. 2344, 2020.
[8] L. Nolte and S. Tomforde, “A Helping Hand: A Survey About AI-Driven Experimental Design for Accelerating Scientific Research,” Applied Sciences, vol. 15, no. 9, p. 5208, 2025.
[9] Z. Zhang, H. Li, S. Jiang, et al., “A survey and evaluation of Web- based tools/databases for variant analysis of TCGA data,” Briefings in Bioinformatics, vol. 20, no. 4, pp. 1524–1541, 2019.
[10] Y. Xu, G. Wu, J. Li, et al., “Screening and Identification of Key Biomarkers for Bladder Cancer: A Study Based on TCGA and GEO Data,” BioMed Research International, vol. 2020, Article ID 8283401, 2020.
[11] S. Kakarmath, A. Esteva, R. Arnaout, et al., “Best practices for authors of healthcare-related artificial intelligence manuscripts,” npj Digital Medicine, vol. 3, no. 1, p. 134, 2020.
[12] C. Meldrum, M. A. Doyle, and R. W. Tothill, “Next-Generation Se- quencing for Cancer Diagnostics: a Practical Perspective,” Clinical Biochemist Reviews, vol. 32, no. 4, pp. 177–195, 2011.
[13] A. M. Hasan, H. A. Jalab, F. Meziane, H. Kahtan, and A. S. Al-Ahmad, “Combining Deep and Handcrafted Image Features for MRI Brain Scan Classification,” IEEE Access, vol. 7, pp. 79959–79967, 2019.
[14] B. Smith, M. Hermsen, E. Lesser, D. Ravichandar, and W. Kremers, “Developing image analysis pipelines of whole-slide images: Pre- and post-processing,” Journal of Clinical and Translational Science, vol. 5, p. e38, 2020.
[15] A. Englisz, M. Smycz-Kuban´ska, and A. Mielczarek-Palacz, “Sensitivity and Specificity of Selected Biomarkers and Their Combinations in the Diagnosis of Ovarian Cancer,” Diagnostics, vol. 14, no. 9, p. 949, 2024.
[16] K. Shah, K. Leow, A. Janssen, T. Shaw, C. Stewart, and I. Kerridge, “Ethical and legal considerations governing use of health data for quality improvement and performance management: a scoping review of the perspectives of health professionals and administrators,” BMJ Open Quality, vol. 14, p. e003309, 2025.
[17] Y. Zhao, R. Gulati, J. Lange, et al., “Sensitivity Measures in Studies of Cancer Early Detection Biomarkers,” Supplementary Materials and Methods, pp. 1–5.
[18] D. Gao, K. Li, R. Wang, S. Shan, and X. Chen, “Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene Text,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019, pp. 12746–12756.
[19] S. M. Raea, K. M. Almotairi, A. M. Alharbi, et al., “Ethical consider- ations in the use of patient medical records for research,” International Journal of Health Sciences, vol. 7, no. S1, pp. 3829–3841, 2023.
[20] C. O. Dumitru and M. Datcu, “Information Content of Very High Resolution SAR Images: Study of Feature Extraction and Imaging Parameters,” IEEE Transactions on Geoscience and Remote Sensing, vol. 51, no. 8, pp. 4591–4610, 2013.
[21] V. W. Lumumba, D. Kiprotich, M. L. Mpaine, N. G. Makena, and M. D. Kavita, “Comparative Analysis of Cross-Validation Tech- niques: LOOCV, K-folds Cross-Validation, and Repeated K-folds Cross- Validation in Machine Learning Models,” American Journal of Theoret- ical and Applied Statistics, vol. 13, no. 5, pp. 127–137, 2024.
[22] M. Ennab and H. Mcheick, “Advancing AI Interpretability in Medical Imaging: A Comparative Analysis of Pixel-Level Interpretability and Grad-CAM Models,” Machine Learning and Knowledge Extraction, vol. 7, no. 1, p. 12, 2025.
[23] R. Vuokko, A. Vakkuri, and S. Palojoki, “Systematized Nomenclature of Medicine–Clinical Terminology (SNOMED CT) Clinical Use Cases in the Context of Electronic Health Record Systems: Systematic Literature Review,” JMIR Medical Informatics, vol. 11, p. e43750, 2023.
[24] M. Mann, C. Kumar, W.-F. Zeng, and M. T. Strauss, “Artificial intelli- gence for proteomics and biomarker discovery,” Cell Systems, vol. 12, pp. 759–770, 2021.
[25] M. Adnan, S. Kalra, J. C. Cresswell, G. W. Taylor, and H. R. Tizhoosh, “Federated learning and differential privacy for medical image analysis,” Scientific Reports, vol. 12, no. 1, p. 1953, 2022.
[26] Y. Chen, et al., “UMPSNet: A Unified Model for Multi-cancer Prog- nostic Survey across Multiple Pathological Slides,” arXiv preprint arXiv:2401.07016, 2024.
[27] S. Rasool, “Integrative Relational Learning on Multimodal Oncology Data,” Moffitt Cancer Center Research, 2024.
[28] Y. Xu, et al., “MUFASA: Multimodal Fusion Architecture Search for Electronic Health Records,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 12, pp. 10532–10540, 2021.
[29] A. Sharma, et al., “Systematic Review of Hybrid Vision Transformer Architectures for Radiological Image Analysis,” Journal of Imaging Informatics in Medicine, 2025.
[30] M. Maillard, et al., “KD-Net: A Knowledge Distillation framework for multi-modal to mono-modal segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, 2020, pp. 38–47.
[31] J. Chen, et al., “Fair Machine Learning in Healthcare: A Review,” ACM Computing Surveys, vol. 55, no. 1, pp. 1–38, 2022.
[32] S. Pfohl, et al., “On the fairness of machine learning in healthcare: dataset shifts and mitigation,” Nature Communications, vol. 14, no. 1, p. 7093, 2023.
How to cite this paper
@article{1711765,
author = {Ayush Mishra, Anadi Mishra, Adarsh Tiwari, Uttam Sharma, Nikhil Raj},
title = {A Unified Multi-Modal Transformer Framework for Synergistic Cancer Diagnosis},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {5},
pages = {197-205},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1711765.pdf},
abstract = {Early cancer diagnosis is critical for improving patient outcomes but is challenged by the disease?s profound heterogeneity. This paper introduces a unified, AI-powered framework that synergistically integrates histopathology, genomics, and proteomics data to enhance early cancer detection. Our architecture features a novel multi-transformer model with dedicated Vision and Genomic Transformers to encode modality-specific features, which are then fused by a cross-modal attention transformer. This intermediate fusion strategy enables the model to learn intricate genotype-phenotype correlations often missed by traditional methods. Validated on cohorts from The Cancer Genome Atlas (TCGA), our framework demonstrates a significant improvement in diagnostic performance over single-modality baselines. We also incorporate Explainable AI (XAI) techniques to ensure model transparency, a crucial step for clinical adoption. The framework serves as both a powerful diagnostic tool and a hypothesis- generation engine, uncovering novel biomarkers from complex multi-modal data and advancing computational pathology and personalized medicine.},
keywords = {Multi-Modal Learning, Transformers, Histopathology, Genomics, Proteomics, Explainable AI, Early Cancer Diagnosis},
month = {November},
}