Home / Current Issue / Paper 1714721
A Multi-Stage Deep Learning Framework for Document Image Restoration
Subject area: Science,Engineering and Technology · Area of research: Computer Vision -> Document Enhancement
Abstract
Real-world camera-captured document images often exhibit complex degradations, including cast shadows, non-uniform illumination, and contrast distortion, which severely degrade visual quality and prevent robust document analysis. In this paper, we present an illumination estimation multi-stage deep learning framework for restoring document images that explicitly separates shadow suppression from illumination normalization. The proposed pipeline involves an initial deep network estimating and mitigating shadow-induced intensity variations before a refinement network corrects global illumination consistency while maintaining textual structure and fine document details. By decomposing the enhancement task into complementary stages, the framework effectively copes with both local shadow artifacts and global lighting imbalance in unconstrained document imaging scenarios. Extensive experiments on real-world camera-captured document images reveal that the proposed method provides visually coherent enhancement with more readable results compared to conventional image processing techniques and existing deep learning-based methods. Standard image quality metrics have been quantitatively evaluated, showing notable gains. The results indicate that the proposed framework offers a robust and practical preprocessing solution for analyzing camera-based document images.
Keywords
Document image enhancement, Shadows and Illumination, Multi-Stage Deep Learning, Camera-Captured Documents
References
[1] J. Zhang, C. Liu, J. Yu, and Z. Guo, “Appearance Enhancement for Camera-Captured Document Images in the Wild,” IEEE Transactions on Artificial Intelligence, vol. 5, no. 5, pp. 2314–2328, May 2024.
[2] A. Das, H. Ma, and S. Kar, “DocProj: A Projective Distortion Dataset for Document Image Enhancement,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1474–1483, 2020.
[3] X. Yi, Y. Zhou, and L. He, “Doc3D: A Large-Scale Synthetic Dataset for Document Image Deformation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2252–2260, 2016.
[4] X. Yi, L. Zhang, and J. Yu, “Doc3DShade: A Synthetic Dataset for Shadow Removal in Document Images,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 157–165, 2017.
[5] J. Jung, S. Cho, and N. I. Cho, “Illumination Correction for Document Images Using Surface Reconstruction,” IEEE Transactions on Image Processing, vol. 26, no. 9, pp. 4362–4376, Sept. 2017.
[6] R. Hedjam, R. F. Moghaddam, and M. Cheriet, “A Statistical Approach for Illumination Compensation of Historical Documents,” IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 5111–5121, Dec. 2012.
[7] Z. Liu et al., “UNeXt: MLP-Based Rapid Medical Image Segmentation Network,” Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2022.
[8] A. Mittal, A. K. Moorthy, and A. C. Bovik, “No-Reference Image Quality Assessment in the Spatial Domain,” IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 4695–4708, Dec. 2012.
[9] R. Smith, “An Overview of the Tesseract OCR Engine,” Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), pp. 629–633, 2007.
[10] A. Bissacco, M. Cummins, Y. Netzer, and H. Neven, “Photo OCR: Reading Text in Uncontrolled Conditions,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 785–792, 2013.
[11] L. Kang, Y. Li, and D. Doermann, “Orientation Robust Text Line Detection in Natural Images,” IEEE Transactions on Image Processing, vol. 25, no. 9, pp. 4182–4195, Sept. 2016.
[12] S. Lu, B. Su, and C. L. Tan, “Document Image Binarization Using Background Estimation and Stroke Edges,” International Journal on Document Analysis and Recognition (IJDAR), vol. 13, no. 4, pp. 303–314, 2010.
[13] S. Xie, Z. Tu, “Holistically-Nested Edge Detection,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 1395–1403, 2015.
[14] C. Tensmeyer and T. Martinez, “Document Image Binarization with Fully Convolutional Neural Networks,” ICDAR, pp. 99–104, 2017.
[15] Y. Wang, J. Zhang, and Z. Guo, “Document Image Appearance Enhancement via Deep Learning,” Pattern Recognition, vol. 95, pp. 292–304, 2019.
[16] D. Karatzas et al., “ICDAR 2015 Competition on Robust Reading,” ICDAR, pp. 1156–1160, 2015.
How to cite this paper
@article{1714721,
author = {Inukollu Anantha Prakash Reddy, Veesam Venkata Srinivas, Gosu Madhu, Dharmavarapu Jayaraju, Gadipudi Krishna Vamsi},
title = {A Multi-Stage Deep Learning Framework for Document Image Restoration},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {8},
pages = {2080-2088},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1714721.pdf},
abstract = {Real-world camera-captured document images often exhibit complex degradations, including cast shadows, non-uniform illumination, and contrast distortion, which severely degrade visual quality and prevent robust document analysis. In this paper, we present an illumination estimation multi-stage deep learning framework for restoring document images that explicitly separates shadow suppression from illumination normalization. The proposed pipeline involves an initial deep network estimating and mitigating shadow-induced intensity variations before a refinement network corrects global illumination consistency while maintaining textual structure and fine document details. By decomposing the enhancement task into complementary stages, the framework effectively copes with both local shadow artifacts and global lighting imbalance in unconstrained document imaging scenarios. Extensive experiments on real-world camera-captured document images reveal that the proposed method provides visually coherent enhancement with more readable results compared to conventional image processing techniques and existing deep learning-based methods. Standard image quality metrics have been quantitatively evaluated, showing notable gains. The results indicate that the proposed framework offers a robust and practical preprocessing solution for analyzing camera-based document images.},
keywords = {Document image enhancement, Shadows and Illumination, Multi-Stage Deep Learning, Camera-Captured Documents},
month = {February},
doi = {https://doi.org/10.64388/IREV9I8-1714721}
}