Home / Current Issue / Paper 1719861
AI-Powered Sign Language Recognition: A Comprehensive Review of CNN, LSTM, Transformers, and MediaPipe Approaches
Subject area: Science,Engineering and Technology · Area of research: Sign Language Recognition
Abstract
Sign language recognition represents a critical application domain for advancing human-computer interaction and accessibility technologies. This paper presents a comprehensive review of deep learning approaches for real-time sign language recognition, with emphasis on convolutional neural networks (CNNs), long short-term memory networks (LSTMs), transformer architectures, and the MediaPipe framework. We analyze the strengths and limitations of each approach, discuss their practical implementations, and examine recent breakthroughs in pose estimation and gesture recognition. Through systematic comparison of state-of-the-art methods, we identify key challenges in achieving robust, real-time performance across diverse sign language dialects and environmental conditions. Furthermore, we explore emerging applications in accessibility, education, and communication systems. This survey provides researchers and practitioners with actionable insights into selecting appropriate architectures for sign language recognition tasks and highlights future research directions in this rapidly evolving field.
Keywords
Sign Language Recognition, Convolutional Neural Networks, Recurrent Neural Networks, Transformers, Mediapipe, Pose Estimation, Deep Learning, Real-Time Processing
References
[1] T. N. Sainath et al., "Deep convolutional neural networks for large-vocabulary continuous speech recognition," in Proc. Interspeech, Portland, OR, USA, Sep. 2012, pp. 338–341.
[2] Z. Huang, X. Long, C. Fang, S. Wang, and M. Qi, "Gesture recognition using multi-stream positional CNN with sequence level supervision," IEEE Trans. Circuits Syst. Video Technol., vol. 31, no. 7, pp. 2786–2797, Jul. 2021.
[3] L. Li, M. Zhou, X. Zhu, and Z. Deng, "Isolated sign language recognition with Convolutional Recurrent Neural Networks," in Proc. Int. Conf. Multimedia Expo (ICME), Hong Kong, Jul. 2017, pp. 745–750.
[4] J. Carreira and A. Zisserman, "Quo vadis, action recognition? A new model and large-scale datasets," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV, USA, Jun. 2016, pp. 6299–6308.
[5] X. Pu, L. Gao, Z. Song, X. Wu, and X. Hong, "Fingerspelling recognition with temporal convolutional networks," in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Seoul, Korea (South), Oct. 2019, pp. 10089–10098.
[6] A. Dosovitskiy et al., "An image is worth 16×16 words: Transformers for image recognition at scale," in Proc. Int. Conf. Learn. Represent. (ICLR), Virtual, May 2021.
[7] Y. Liu, K. Lin, Y. Li, Y. Wang, C. Gao, and W. Yu, "Sign language recognition with transformers," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), New Orleans, LA, USA, Jun. 2022, pp. 1556–1565.
[8] F. Zhang, V. Bazarevsky, A. Vakunov, A. Tkacenko, and M. Sung, "Mediapipe hands: On-device real-time hand tracking," arXiv preprint arXiv:2006.10214, Jun. 2020.
[9] K. Saunders, J. Camgöz, and R. Bowden, "Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Los Angeles, CA, USA, Jun. 2020, pp. 7210–7219.
[10] H. Zhou, W. Zhou, Y. Zhou, and R. Mao, "Deep spatial-temporal convolution neural networks for sign language recognition," in Proc. AAAI Conf. Artif. Intell., vol. 35, no. 4, pp. 3610–3617, 2021.
[11] M. Joze and O. Koller, "MS-ASL: A large-scale dataset for American sign language recognition," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Los Angeles, CA, USA, Jun. 2020, pp. 10092–10101.
[12] S. Vaswani et al., "Attention is all you need," in Proc. 31st Conf. Neural Inf. Process. Syst. (NeurIPS), Long Beach, CA, USA, Dec. 2017, pp. 5998–6008.
How to cite this paper
@article{1719861,
author = {Udit Mehla, Amit Chobey, Vinay Kumar, Lucky, Monika},
title = {AI-Powered Sign Language Recognition: A Comprehensive Review of CNN, LSTM, Transformers, and MediaPipe Approaches},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {7},
number = {7},
pages = {834-838},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1719861.pdf},
abstract = {Sign language recognition represents a critical application domain for advancing human-computer interaction and accessibility technologies. This paper presents a comprehensive review of deep learning approaches for real-time sign language recognition, with emphasis on convolutional neural networks (CNNs), long short-term memory networks (LSTMs), transformer architectures, and the MediaPipe framework. We analyze the strengths and limitations of each approach, discuss their practical implementations, and examine recent breakthroughs in pose estimation and gesture recognition. Through systematic comparison of state-of-the-art methods, we identify key challenges in achieving robust, real-time performance across diverse sign language dialects and environmental conditions. Furthermore, we explore emerging applications in accessibility, education, and communication systems. This survey provides researchers and practitioners with actionable insights into selecting appropriate architectures for sign language recognition tasks and highlights future research directions in this rapidly evolving field.},
keywords = {Sign Language Recognition, Convolutional Neural Networks, Recurrent Neural Networks, Transformers, Mediapipe, Pose Estimation, Deep Learning, Real-Time Processing},
month = {January},
}