International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1719448

1719448 Vol 9 · Issue 12 Download Paper

Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition

Saiba Teja Sri Sanaboina Chandra Sekhar

Subject area: Science,Engineering and Technology  ·  Area of research: Deep Learning

DOI: 10.64388/IREV9I12-1719448

Abstract

Hand gesture recognition plays significant role in HCI applications and also it has been incorporated in the fields of Virtual Reality (VR), Augmented Reality (AR) and Robotics. Gestures in the real world are continuous and under these conditions, any recognition online would be very difficult since there are two types of gesture classes to be recognized as well as the temporal limits of these classes of gesture. Current state-of-the-art deep learning based on sliding windows and attention mechanisms yield piece-wise and “jagged” predictions resulting in more false alarms as well as incorrect gesture boundaries. Single-modality methods are not beneficial for recognition purposes because they do not return full-fledged gestures of the gesture. It is a reproduction of a cross-attention-based online hand gesture recognition model based on Joint Collection Distance (JCD) and Frame Vector (FV) that reproduces attention-based models. Taking the temporal refinement approach is proposed to reduce the noise and enhance the detection of the boundaries. To extend this into multimodal, visual appearance features of video frames are added in through an existing CNN which can learn structural and appearance features. Experimental findings of IPN Hand Gesture Dataset indicate that the Temporal refinement enhanced the Detection Rate of 0.9184 to 0.9288, False Positives of 39 to 14, and Mean IoU of 0.7375 to 0.8153 which showed that the gesture recognition and temporal localization were more accurate.

Keywords

Hand Gesture Recognition, Multimodal Learning, Attention Mechanism, Temporal Refinement, Joint Collection Distance (JCD), Frame Vector (FV), Deep Learning, Sliding Window.

References

[1] M. J. Chae, S. H. Han, H. Nam, J. H. Park, M. H. Cha, and S. I. Cho, "Online Hand Gesture Recognition Using Semantically Interpretable Attention Mechanism," IEEE Access, vol. 13, pp. 32329–32340, 2025, doi: 10.1109/ACCESS.2025.3540721.

[2] G. Benitez-Garcia, J. Olivares-Mercado, G. Sanchez-Perez, and K. Yanai, "IPN Hand: A Video Dataset and Benchmark for Real-Time Continuous Hand Gesture Recognition," in Proc. IEEE Int. Conf. Pattern Recognit., 2021.

[3] O. Köpüklü, A. Gunduz, N. Kose, and G. Rigoll, "Online Dynamic Hand Gesture Recognition Including Efficiency Analysis," IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 2, no. 2, pp. 85–97, Apr. 2020, doi: 10.1109/TBIOM.2020.2977750.

[4] P. Molchanov, X. Yang, S. Gupta, K. Kim, and J. Kautz, "Online Detection and Classification of Dynamic Hand Gestures with Recurrent 3D Convolutional Neural Networks," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 4207–4215, doi: 10.1109/CVPR.2016.456.

[5] H. Chen, X. Liu, J. Shi, and G. Zhao, "Temporal Hierarchical Dictionary Guided Decoding for Online Gesture Segmentation and Recognition," IEEE Transactions on Image Processing, vol. 29, pp. 9689–9702, 2020, doi: 10.1109/TIP.2020.3028962.

[6] S. Zhu, X. Pan, X. Cheng, and H. Guo, "Continuous Gesture Segmentation and Recognition Using 3DCNN and Convolutional LSTM," IEEE Transactions on Multimedia, vol. 21, no. 4, pp. 1011–1021, Apr. 2019, doi: 10.1109/TMM.2018.2869278.

[7] C. Lea, M. Flynn, R. Vidal, A. Reiter, and G. Hager, "Temporal Convolutional Networks for Action Segmentation and Detection," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2017, pp. 156–165, doi: 10.1109/CVPR.2017.113.

[8] X. Yang, P. Molchanov, and J. Kautz, "Making Convolutional Networks Recurrent for Visual Sequence Learning," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2018, pp. 6469–6478, doi: 10.1109/CVPR.2018.00677.

[9] H. Wang, P. Wang, Z. Song, and W. Li, "Large-Scale Multimodal Gesture Recognition Using Heterogeneous Networks," IEEE Transactions on Cybernetics, vol. 49, no. 10, pp. 3660–3673, Oct. 2019, doi: 10.1109/TCYB.2018.2844809.

[10] M. Abavisani, H. Joze, and V. M. Patel, "Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 1165–1174, doi: 10.1109/CVPR42600.2020.00124.

[11] R. Kopuklu, N. Kose, A. Gunduz, and G. Rigoll, "Real-time Hand Gesture Detection and Classification Using Convolutional Neural Networks," in Proc. IEEE Int. Conf. Autom. Face Gesture Recognit., 2019, doi: 10.1109/FG.2019.8756576.

[12] L. Qu, H. Wu, T. Yang, L. Zhang, and Y. Sun, "Dynamic Hand Gesture Classification Based on Multichannel Radar Using Multistream Fusion 1-D Convolutional Neural Network," IEEE Sensors Journal, vol. 22, no. 24, pp. 24083–24093, Dec. 2022, doi: 10.1109/JSEN.2022.3216604.

[13] W. Zhang, J. Wang, and F. Lan, "Dynamic hand gesture recognition based on short-term sampling neural networks," IEEE/CAA Journal of Automatica Sinica, vol. 8, no. 1, pp. 110–120, Jan. 2021, doi: 10.1109/JAS.2020.1003465.

[14] D. Zhao, H. Li, and S. Yan, "Spatial-Temporal Synchronous Transformer for Skeleton-Based Hand Gesture Recognition," IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 3, pp. 1403–1412, Mar. 2024, doi: 10.1109/TCSVT.2023.3295084.

[15] Y. Cao, J. Li, C. Chakraborty, L. Qin, L. Tao, and X. Shao, "Temporal Segment Neural Networks-Enabled Dynamic Hand-Gesture Recognition for Industrial Cyber-Physical Authentication Systems," IEEE Systems Journal, vol. 17, no. 4, pp. 5315–5326, Dec. 2023, doi: 10.1109/JSYST.2023.3306380.

[16] C. Si, W. Chen, W. Wang, L. Wang, and T. Tan, "An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action Recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2019, doi: 10.1109/CVPR.2019.00132.

[17] L. Shi, Y. Zhang, J. Cheng, and H. Lu, "Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2019, doi: 10.1109/CVPR.2019.00054.

[18] L. Shi, Y. Zhang, J. Cheng, and H. Lu, "Skeleton-Based Action Recognition with Directed Graph Neural Networks," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2019, doi: 10.1109/CVPR.2019.00810.

[19] Y. Du, W. Wang, and L. Wang, "Hierarchical Recurrent Neural Network for Skeleton Based Action Recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2015, pp. 1110–1118, doi: 10.1109/CVPR.2015.7298714.

[20] H. Wang and L. Wang, "Modeling Temporal Dynamics and Spatial Configurations of Actions Using Two-Stream Recurrent Neural Networks," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2017, pp. 499–508, doi: 10.1109/CVPR.2017.61.

[21] Z. Liu, H. Zhang, Z. Chen, Z. Wang, and W. Ouyang, "Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 143–152, doi: 10.1109/CVPR42600.2020.00022.

[22] X. Chen, Y. Ye, and J. Xu, "Richly Activated Graph Convolutional Network for Robust Skeleton-Based Action Recognition," IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 5, pp. 1915–1925, May 2021, doi: 10.1109/TCSVT.2020.3015051.

[23] Z. Tu, H. Zhang, H. Liu, J. Yuan, and J. Li, "Constructing Stronger and Faster Baselines for Skeleton-Based Action Recognition," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 1474–1488, Feb. 2023, doi: 10.1109/TPAMI.2022.3157033.

[24] J. Liu, A. Shahroudy, M. Perez, G. Wang, L. Duan, and A. C. Kot, "NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 10, pp. 2684–2701, Oct. 2020, doi: 10.1109/TPAMI.2019.2916873.

[25] A. Shahroudy, J. Liu, T. Ng, and G. Wang, "NTU RGB+D: A Large-Scale Dataset for 3D Human Activity Analysis," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 1010–1019, doi: 10.1109/CVPR.2016.115.

[26] J. Y. Kim, B. H. Kang, and D. H. Im, "Gate-Shift Networks for Video Action Recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, doi: 10.1109/CVPR42600.2020.01245.

[27] D. Tran et al., "Learning Spatiotemporal Features With 3D Convolutional Networks," in Proc. IEEE Int. Conf. Comput. Vis., 2015, pp. 4489–4497, doi: 10.1109/ICCV.2015.510.

[28] T. Zhang, W. Zheng, Z. Cui, C. Shan, and J. Yang, "A Deep Neural Framework for Continuous Sign Language Recognition by Iterative Training," IEEE Transactions on Multimedia, vol. 21, no. 7, pp. 1880–1891, Jul. 2019, doi: 10.1109/TMM.2018.2889563.

[29] J. Alon, V. Athitsos, Q. Yuan, and S. Sclaroff, "A unified framework for gesture recognition and spatiotemporal gesture segmentation," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 31, no. 9, pp. 1685–1699, 2009, doi: 10.1109/TPAMI.2008.203.

[30] A. Osman Hashi, S. Zaiton Mohd Hashim, and A. Bte Asamah, "A Systematic Review of Hand Gesture Recognition: An Update from 2018 to 2024," IEEE Access, 2024, doi: 10.1109/ACCESS.2024.3421992.

[31] Z. R. Saeed, Z. B. Zainol, B. B. Zaidan, and A. H. Alamoodi, "A Systematic Review on Systems-Based Sensory Gloves for Sign Language Pattern Recognition: An Update from 2017 to 2022," IEEE Access, 2022, doi: 10.1109/ACCESS.2022.3219430.

How to cite this paper

Saiba Teja Sri, Sanaboina Chandra Sekhar "Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition" Iconic Research And Engineering Journals Volume 9 Issue 12 2026 Page 3661-3674 https://doi.org/10.64388/IREV9I12-1719448
Saiba Teja Sri, Sanaboina Chandra Sekhar "Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition" Iconic Research And Engineering Journals, vol. 9, no. 12, Jun. 2026, doi: https://doi.org/10.64388/IREV9I12-1719448
Saiba Teja Sri, Sanaboina Chandra Sekhar (2026). Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition. Iconic Research And Engineering Journals, 9(12). doi: https://doi.org/10.64388/IREV9I12-1719448
Saiba Teja Sri, Sanaboina Chandra Sekhar "Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition" Iconic Research And Engineering Journals, vol. 9, no. 12, Jun. 2026. Crossref, https://doi.org/10.64388/IREV9I12-1719448
@article{1719448,
      author = {Saiba Teja Sri, Sanaboina Chandra Sekhar},
      title = {Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {12},
      pages = {3661-3674},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1719448.pdf},
      abstract = {Hand gesture recognition plays significant role in HCI applications and also it has been incorporated in the fields of Virtual Reality (VR), Augmented Reality (AR) and Robotics. Gestures in the real world are continuous and under these conditions, any recognition online would be very difficult since there are two types of gesture classes to be recognized as well as the temporal limits of these classes of gesture. Current state-of-the-art deep learning based on sliding windows and attention mechanisms yield piece-wise and “jagged” predictions resulting in more false alarms as well as incorrect gesture boundaries. Single-modality methods are not beneficial for recognition purposes because they do not return full-fledged gestures of the gesture. It is a reproduction of a cross-attention-based online hand gesture recognition model based on Joint Collection Distance (JCD) and Frame Vector (FV) that reproduces attention-based models. Taking the temporal refinement approach is proposed to reduce the noise and enhance the detection of the boundaries. To extend this into multimodal, visual appearance features of video frames are added in through an existing CNN which can learn structural and appearance features. Experimental findings of IPN Hand Gesture Dataset indicate that the Temporal refinement enhanced the Detection Rate of 0.9184 to 0.9288, False Positives of 39 to 14, and Mean IoU of 0.7375 to 0.8153 which showed that the gesture recognition and temporal localization were more accurate.},
      keywords = {Hand Gesture Recognition, Multimodal Learning, Attention Mechanism, Temporal Refinement, Joint Collection Distance (JCD), Frame Vector (FV), Deep Learning, Sliding Window.},
      month = {June},
      doi = {https://doi.org/10.64388/IREV9I12-1719448}
  }