Home / Current Issue / Paper 1719861
AI-Powered Sign Language Recognition: A Comprehensive Review of CNN, LSTM, Transformers, and MediaPipe Approaches
Subject area: Science,Engineering and Technology · Area of research: Sign Language Recognition
Abstract
Sign language recognition represents a critical application domain for advancing human-computer interaction and accessibility technologies. This paper presents a comprehensive review of deep learning approaches for real-time sign language recognition, with emphasis on convolutional neural networks (CNNs), long short-term memory networks (LSTMs), transformer architectures, and the MediaPipe framework. We analyze the strengths and limitations of each approach, discuss their practical implementations, and examine recent breakthroughs in pose estimation and gesture recognition. Through systematic comparison of state-of-the-art methods, we identify key challenges in achieving robust, real-time performance across diverse sign language dialects and environmental conditions. Furthermore, we explore emerging applications in accessibility, education, and communication systems. This survey provides researchers and practitioners with actionable insights into selecting appropriate architectures for sign language recognition tasks and highlights future research directions in this rapidly evolving field.
Keywords
Sign Language Recognition, Convolutional Neural Networks, Recurrent Neural Networks, Transformers, Mediapipe, Pose Estimation, Deep Learning, Real-Time Processing
How to cite this paper
@article{1719861,
author = {Udit Mehla, Amit Chobey, Vinay Kumar, Lucky, Monika},
title = {AI-Powered Sign Language Recognition: A Comprehensive Review of CNN, LSTM, Transformers, and MediaPipe Approaches},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {7},
number = {7},
pages = {834-838},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1719861.pdf},
abstract = {Sign language recognition represents a critical application domain for advancing human-computer interaction and accessibility technologies. This paper presents a comprehensive review of deep learning approaches for real-time sign language recognition, with emphasis on convolutional neural networks (CNNs), long short-term memory networks (LSTMs), transformer architectures, and the MediaPipe framework. We analyze the strengths and limitations of each approach, discuss their practical implementations, and examine recent breakthroughs in pose estimation and gesture recognition. Through systematic comparison of state-of-the-art methods, we identify key challenges in achieving robust, real-time performance across diverse sign language dialects and environmental conditions. Furthermore, we explore emerging applications in accessibility, education, and communication systems. This survey provides researchers and practitioners with actionable insights into selecting appropriate architectures for sign language recognition tasks and highlights future research directions in this rapidly evolving field.},
keywords = {Sign Language Recognition, Convolutional Neural Networks, Recurrent Neural Networks, Transformers, Mediapipe, Pose Estimation, Deep Learning, Real-Time Processing},
month = {January},
}