Home / Current Issue / Paper 1716845
Multimodal Sensor-Agnostic Gesture- Controlled Gaming Interface with Adaptive Depth Integration
Subject area: Science,Engineering and Technology · Area of research: Computer Science
DOI: https://doi.org/10.64388/IREV9I10-1716845
Abstract
Despite the fact that HCI through gestures has gained significant popularity due to their intuitiveness in recent years, the current gesture recognition systems lack robustness since they rely on hardware support, use a single modality as input, and have few practical applications in real-world settings. In this paper, a gesture recognition algorithm that does not require any hardware support except for the camera sensor is introduced, taking into account RGB data and Microsoft Kinect depth sensors. In particular, the gesture recognizer is based on the detection of the joints of a skeleton and hand landmarks. The mathematical formula used for the proposed multimodal representation is shown below. F_t = [S_t || H_t || D_t]. These experiments are carried out using our own customized gesture dataset that includes 4,000 gestures from 4 categories from 10 different subjects. Recognition accuracy is 91.6% and latency is 34 ms in real-time operation mode, which are significantly better than the outcomes of the baselines based on individual modality.
Keywords
Gesture Recognition, Human Computer Interaction, Multimodal Fusion, LSTM, MediaPipe, Depth Sensing
References
[1] J. Shotton, A. Fitzgibbon, M. Cook, T. Sharp, M. Finocchio, R. Moore, A. Kipman, and A. Blake, “Real- Time Human Pose Recognition in Parts from Single Depth Images,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2011, pp. 1297–1304.
[2] C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Ubowejeke, M. Hays, F. Zhang, C. Chang, M. Wan, and T. Grundmann, “MediaPipe: A Framework for Building Perception Pipelines,” arXiv preprint arXiv:1906.08172, 2019.
[3] P. Wang, W. Li, P. Ogunbona, J. Wan, and S. Escalera, “RGB-D-Based Human Motion Recognition with Deep Learning: A Survey,” Comput. Vis. Image Underst., vol. 171, pp. 118–137, 2018.
[4] X. Zhang, X. Liu, J. Yuan, and S. Lin, “Hand Gesture Recognition Based on Deep Learning: A Review,” IEEE Access, vol. 8, pp. 45 765– 45 777, 2020.
[5] N. Neverova, C. Wolf, G. Taylor, and F. Nebout, “ModDrop: Adaptive Multi-Modal Gesture Recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 38, no. 8, pp. 1692– 1706, 2016.
[6] Z. Ouyang, J. Cui, and S. Liu, “Multimodal Human Activity Recognition: A Review and New Perspectives,” Inf. Fusion, vol. 56, pp. 116–143, 2020.
[7] J. Donahue, L. A. Hendricks, M. Rohrbach, S. Venugopalan, S. Guadarrama, K. Saenko, and T. Darrell, “Long-Term Recurrent Convolutional Networks for Visual Recognition and Description,” in Proc. IEEE CVPR, 2015, pp. 2625–2634.
[8] Y. Du, W. Wang, and L. Wang, “Hierarchical Recurrent Neural Network for Skeleton Based Action Recognition,” in Proc. IEEE CVPR, 2015, pp. 1110–1118.
[9] S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Comput., vol. 9, no. 8, pp. 1735–1780,1997.
[10] F. Zhang, V. Bazarevsky, A. Vakunov, A. Tkachenka, G. Sung, C.-L. Chang, and M. Grundmann, “MediaPipe Hands: On-device Real-time Hand Tracking,” arXiv preprint arXiv:2006.10214, 2020.
[11] V. Bazarevsky, I. Grishchenko, K. Raveendran, T. Zhu, F. Zhang, and M. Grundmann, “BlazePose: On-Device Real- Time Body Pose Tracking,” arXiv preprint arXiv:2006.10204, 2020.
[12] S. Biswas, A. Nandy, A. K. Naskar, and R. Saw, “MediaPipe with LSTM Architecture for Real-Time Hand Gesture Recognition,” in Proc. CVIP 2023, Commun. Comput. Inf. Sci., vol. 2010, Springer, 2024, pp. 427–438.
[13] R. Rastgoo, K. Kiani, S. Escalera, and M. Sabokrou, “Multi-Modal Zero-Shot Dynamic Hand Gesture Recognition,” Expert Syst. Appl., vol. 247, p. 123349, 2024.
[14] P. Balaji and M. R. Prusty, “Multimodal Fusion Hierarchical Self-Attention Network for Dynamic Hand Gesture Recognition,” J. Vis. Commun. Image Represent., vol. 98, p. 104019, 2024.2023.
[15] N. C. Mithun, N. U. Ahmed, and S. M. M. Rahman, “Activity Recognition Using Fusion of Low-Cost Sensors,” IEEE Trans. Consum. Electron., vol. 63, no. 4,pp. 428–438, 2017.
[16] A. Graves, A. Mohamed, and G. Hinton, “Speech Recognition with Deep Recurrent Neural Networks,” in Proc. IEEE ICASSP, 2013, pp. 6645–6649.
[17] P. Huynh, T. Nguyen, and H. Le, “Proposing Hand Gesture Recognition System Using MediaPipe Holistic and LSTM,” in Proc. IEEE ICCE-Asia,
How to cite this paper
@article{1716845,
author = {Sagar Gangal, Rahul Nelogi, Sachidananda K},
title = {Multimodal Sensor-Agnostic Gesture- Controlled Gaming Interface with Adaptive Depth Integration},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {10},
pages = {3184-3191},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1716845.pdf},
abstract = {Despite the fact that HCI through gestures has gained significant popularity due to their intuitiveness in recent years, the current gesture recognition systems lack robustness since they rely on hardware support, use a single modality as input, and have few practical applications in real-world settings. In this paper, a gesture recognition algorithm that does not require any hardware support except for the camera sensor is introduced, taking into account RGB data and Microsoft Kinect depth sensors. In particular, the gesture recognizer is based on the detection of the joints of a skeleton and hand landmarks. The mathematical formula used for the proposed multimodal representation is shown below. F_t = [S_t || H_t || D_t]. These experiments are carried out using our own customized gesture dataset that includes 4,000 gestures from 4 categories from 10 different subjects. Recognition accuracy is 91.6% and latency is 34 ms in real-time operation mode, which are significantly better than the outcomes of the baselines based on individual modality.},
keywords = {Gesture Recognition, Human Computer Interaction, Multimodal Fusion, LSTM, MediaPipe, Depth Sensing},
month = {April},
doi = {https://doi.org/10.64388/IREV9I10-1716845}
}