Home / Current Issue / Paper 1712718
Multi-Lingual Translation using Image
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
Abstract
Multi-lingual translation using images has emerged as a powerful approach to bridge linguistic barriers in real-time communication. This paper presents a system that automatically extracts text from images and translates it into multiple target languages using a combination of Optical Character Recognition (OCR) and Neural Machine Translation (NMT) models. The proposed framework captures an input image, preprocesses it to enhance text visibility, and applies OCR to accurately detect and extract textual content across diverse scripts. A deep learning?based translation engine then converts the extracted text into user-selected languages while preserving contextual meaning. The system supports multiple languages, including English and various regional Indian languages, enabling seamless cross-lingual understanding. Experimental results demonstrate high accuracy in text detection and translation, even under challenging conditions such as noisy backgrounds, varying fonts, and low illumination. This work contributes to the development of intelligent, user-friendly translation tools suitable for education, tourism, document digitization, and assistive technologies.
References
[1] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 6, pp. 1137–1149, 2017.
[2] X. Zhou et al., “EAST: An Efficient and Accurate Scene Text Detector,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 5551–5560.
[3] Y. Zhang, W. Liu, Z. Wang, and X. Gu, “Scene Text Detection and Recognition: The Deep Learning Era,” IEEE Access, vol. 7, pp. 145158–145180, 2019.
[4] R. Smith, “An Overview of the Tesseract OCR Engine,” in Proc. Int. Conf. Document Analysis and Recognition (ICDAR), 2007, pp. 629–633.
[5] D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2015.
[6] A. Vaswani et al., “Attention Is All You Need,” in Proc. Adv. Neural Inf. Process. Syst. (NIPS), 2017, pp. 5998–6008.
[7] M. Artetxe, G. Labaka, and E. Agirre, “Unsupervised Statistical Machine Translation,” in Proc. EMNLP, 2018, pp. 3632–3642.
[8] Y. Liu et al., “mBART: Multilingual Denoising Pre-training for Neural Machine Translation,” in Proc. ACL, 2020, pp. 3645–3657.
[9] S. Sun, C. Zhang, and C. Guo, “A Review of OCR Techniques for Scene Text Recognition,” IEEE Access, vol. 8, pp. 110852–110872, 2020.
[10] H. Li, P. Wang, and C. Shen, “Towards End-to-End Text Spotting with Convolutional Recurrent Neural Networks,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), 2017, pp. 5238–5246.
[11] K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in Proc. ICLR, 2015.
[12] J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. NAACL-HLT, 2019.
[13] A. Conneau et al., “XNLI: Evaluating Cross-lingual Sentence Representations,” in Proc. EMNLP, 2018.
[14] T. Kudo and J. Richardson, “SentencePiece: A Simple and Language Independent Subword Tokenization Algorithm for Neural Text Processing,” in Proc. EMNLP, 2018.
[15] P. Gupta, S. Shetty, and R. Ranjan, “Image-Based Text Extraction and Translation for Indian Regional Languages,” Int. J. Comput. Appl., vol. 178, no. 23, pp. 1–5, 2019.
[16] Google Research, “OCR and Multilingual Translation Technologies,” Google AI Blog, 2020. [Online]. Available: https://ai.googleblog.com.
How to cite this paper
@article{1712718,
author = {R Lohith, Sangamesh D Chidri, Yashwanth K, Ganeshan},
title = {Multi-Lingual Translation using Image},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {6},
pages = {432-439},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1712718.pdf},
abstract = {Multi-lingual translation using images has emerged as a powerful approach to bridge linguistic barriers in real-time communication. This paper presents a system that automatically extracts text from images and translates it into multiple target languages using a combination of Optical Character Recognition (OCR) and Neural Machine Translation (NMT) models. The proposed framework captures an input image, preprocesses it to enhance text visibility, and applies OCR to accurately detect and extract textual content across diverse scripts. A deep learning?based translation engine then converts the extracted text into user-selected languages while preserving contextual meaning. The system supports multiple languages, including English and various regional Indian languages, enabling seamless cross-lingual understanding. Experimental results demonstrate high accuracy in text detection and translation, even under challenging conditions such as noisy backgrounds, varying fonts, and low illumination. This work contributes to the development of intelligent, user-friendly translation tools suitable for education, tourism, document digitization, and assistive technologies.},
month = {December},
doi = {https://doi.org/10.64388/IREV9I6-1712718}
}