International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1718689

1718689 Vol 9 · Issue 12 Download Paper

Comparative Study of Deep Learning ArchitecturesforAutomated Diabetic Retinopathy Grading:Vision Transformer, Swin Transformer, and InceptionResNetV2

Samir Mulla Mahamadtohid Naikwadi Prajwal Khandait Aditya Sutar Rajesh Kumar Uma Gurav

Subject area: Science,Engineering and Technology  ·  Area of research: Artificial Intelligence, Deep learning

DOI: 10.64388/IREV9I12-1718689

Abstract

Diabetic Retinopathy (DR) is a vision-threatening complication of diabetes mellitus that progresses silently through five clinically defined severity grades. Timely automated screen-ing is critical to prevent irreversible vision loss, particu-larly in resource-constrained healthcare settings. This paper presents a systematic comparative study of three state-of-the-art deep learning architectures-Vision Transformer (ViT-Base/16), Swin Transformer (swin base patch4 window7 224), and InceptionResNetV2-applied to five-class DR grading on the APTOS 2019 fundus image dataset (3,662 images). All models employ transfer learning from ImageNet-pretrained weights. We analyze each architecture from the perspectives of classification accuracy, per-class F1-score, macro-averaged AUC, GradCAM-based explainability, training dynamics, and parameter efficiency. Our ViT-Base/16 model, fine-tuned end-to-end with AdamW, cosine annealing, and label smoothing, achieves the highest validation accuracy of 85.40% with a macro-averaged F1-score of 0.7247. Swin Transformer achieves 83.20% accuracy, while InceptionResNetV2 achieves 81.40% through two-stage transfer learning. GradCAM visualizations confirm clinically aligned lesion localization across all architectures. This work provides architectural insights for deploying robust DR screening systems in clinical environments.

Keywords

Diabetic Retinopathy, Vision Transformer, Swin Transformer, InceptionResNetV2, Transfer Learning, GradCAM, Fundus Image Classification, Deep Learning, Medical Image Analysis

References

[1] International Diabetes Federation, “IDF Diabetes Atlas, 10th edition,” 2021. [Online]. Available: https://www.diabetesatlas.org

[2] World Health Organization, “Blindness and vision impairment,” WHO Fact Sheet, 2023. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/blindness-and-visual-impairment

[3] C. P. Wilkinson et al., “Proposed international clinical diabetic retinopa-thy and diabetic macular edema disease severity scales,” Ophthalmology, vol. 110, no. 9, pp. 1677-1682, 2003.

[4] D. S. W. Ting et al., “Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes,” JAMA, vol. 318, no. 22, pp. 2211-2223, 2017.

[5] M. D. Abra`moff et al., “Improved automated detection of diabetic retinopathy on a publicly available dataset through integration of deep learning,” Investigative Ophthalmology & Visual Science, vol. 57, no. 13, pp. 5200-5206, 2016.

[6] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436-444, 2015.

[7] A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. ICLR, 2021.

[8] Z. Liu et al., “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF ICCV, pp. 10012-10022, 2021.

[9] C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi, “Inception-v4, Inception-ResNet and the impact of residual connections on learning,” in Proc. AAAI, 2017.

[10] V. Gulshan et al., “Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus pho-tographs,” JAMA, vol. 316, no. 22, pp. 2402-2410, 2016.

[11] M. D. Abra`moff et al., “Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices,” NPJ Digital Medicine, vol. 1, no. 1, pp. 1-8, 2018.

[12] R. Gargeya and T. Leng, “Automated identification of diabetic retinopa-thy using deep learning,” Ophthalmology, vol. 124, no. 7, pp. 962-969, 2017.

[13] B. Graham, “Kaggle diabetic retinopathy detection competition report,” University of Warwick, 2015.

[14] N. Sikder, M. S. Masud, A. K. M. B. Hossain, and M. A. Bhuiyan, “Severity classification of diabetic retinopathy using an ensemble learn-ing algorithm through analyzing retinal images,” Symmetry, vol. 13, no. 4, p. 670, 2021.

[15] S. Qummar et al., “A deep learning ensemble approach for diabetic retinopathy detection,” IEEE Access, vol. 7, pp. 150530-150539, 2019.

[16] X. Wang et al., “Self-attention-based CNNs for diabetic retinopathy grading using fundus image,” in Proc. IEEE ISBI, 2021.

[17] X. Sun, J. Xu, and J. Ma, “Vision transformer for diabetic retinopathy grading,” in Proc. Int. Conf. on Medical Image Analysis and Computer-Aided Diagnosis, 2021.

[18] B. Gheflati and H. Rivaz, “Vision transformers for classification of diabetic retinopathy,” in Proc. IEEE EMBC, pp. 1988-1991, 2022.

[19] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. IEEE/CVF ICCV, pp. 618-626, 2017.

[20] A. Chattopadhay, A. Sarkar, P. Howlader, and V. N. Balasubramanian, “Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks,” in Proc. IEEE WACV, pp. 839-847, 2018.

[21] H. Chefer, S. Gur, and L. Wolf, “Transformer interpretability beyond attention visualization,” in Proc. IEEE/CVF CVPR, pp. 782-791, 2021.

[22] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR, 2019.

[23] R. Wightman, “PyTorch Image Models (timm),” GitHub, 2019. [Online]. Available: https://github.com/rwightman/pytorch-image-models

[24] Asia Pacific Tele-Ophthalmology Society, “APTOS 2019 Blind-ness Detection,” Kaggle Competition, 2019. [Online]. Available: https://www.kaggle.com/c/aptos2019-blindness-detection

How to cite this paper

Samir Mulla, Mahamadtohid Naikwadi, Prajwal Khandait, Aditya Sutar, Rajesh Kumar; Uma Gurav "Comparative Study of Deep Learning ArchitecturesforAutomated Diabetic Retinopathy Grading:Vision Transformer, Swin Transformer, and InceptionResNetV2" Iconic Research And Engineering Journals Volume 9 Issue 12 2026 Page 527-536 https://doi.org/10.64388/IREV9I12-1718689
Samir Mulla, Mahamadtohid Naikwadi, Prajwal Khandait, Aditya Sutar, Rajesh Kumar; Uma Gurav "Comparative Study of Deep Learning ArchitecturesforAutomated Diabetic Retinopathy Grading:Vision Transformer, Swin Transformer, and InceptionResNetV2" Iconic Research And Engineering Journals, vol. 9, no. 12, Jun. 2026, doi: https://doi.org/10.64388/IREV9I12-1718689
Samir Mulla, Mahamadtohid Naikwadi, Prajwal Khandait, Aditya Sutar, Rajesh Kumar; Uma Gurav (2026). Comparative Study of Deep Learning ArchitecturesforAutomated Diabetic Retinopathy Grading:Vision Transformer, Swin Transformer, and InceptionResNetV2. Iconic Research And Engineering Journals, 9(12). doi: https://doi.org/10.64388/IREV9I12-1718689
Samir Mulla, Mahamadtohid Naikwadi, Prajwal Khandait, Aditya Sutar, Rajesh Kumar; Uma Gurav "Comparative Study of Deep Learning ArchitecturesforAutomated Diabetic Retinopathy Grading:Vision Transformer, Swin Transformer, and InceptionResNetV2" Iconic Research And Engineering Journals, vol. 9, no. 12, Jun. 2026. Crossref, https://doi.org/10.64388/IREV9I12-1718689
@article{1718689,
      author = {Samir Mulla, Mahamadtohid Naikwadi, Prajwal Khandait, Aditya Sutar, Rajesh Kumar; Uma Gurav},
      title = {Comparative Study of Deep Learning ArchitecturesforAutomated Diabetic Retinopathy Grading:Vision Transformer, Swin Transformer, and InceptionResNetV2},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {12},
      pages = {527-536},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1718689.pdf},
      abstract = {Diabetic Retinopathy (DR) is a vision-threatening complication of diabetes mellitus that progresses silently through five clinically defined severity grades. Timely automated screen-ing is critical to prevent irreversible vision loss, particu-larly in resource-constrained healthcare settings. This paper presents a systematic comparative study of three state-of-the-art deep learning architectures-Vision Transformer (ViT-Base/16), Swin Transformer (swin base patch4 window7 224), and InceptionResNetV2-applied to five-class DR grading on the APTOS 2019 fundus image dataset (3,662 images). All models employ transfer learning from ImageNet-pretrained weights. We analyze each architecture from the perspectives of classification accuracy, per-class F1-score, macro-averaged AUC, GradCAM-based explainability, training dynamics, and parameter efficiency. Our ViT-Base/16 model, fine-tuned end-to-end with AdamW, cosine annealing, and label smoothing, achieves the highest validation accuracy of 85.40% with a macro-averaged F1-score of 0.7247. Swin Transformer achieves 83.20% accuracy, while InceptionResNetV2 achieves 81.40% through two-stage transfer learning. GradCAM visualizations confirm clinically aligned lesion localization across all architectures. This work provides architectural insights for deploying robust DR screening systems in clinical environments.},
      keywords = {Diabetic Retinopathy, Vision Transformer, Swin Transformer, InceptionResNetV2, Transfer Learning, GradCAM, Fundus Image Classification, Deep Learning, Medical Image Analysis},
      month = {June},
      doi = {https://doi.org/10.64388/IREV9I12-1718689}
  }