International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1719448

1719448PublishedVol 9 · Issue 12

Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition

Saiba Teja Sri Sanaboina Chandra Sekhar

Subject area: Science,Engineering and Technology  ·  Area of research: Deep Learning

DOI: https://doi.org/10.64388/IREV9I12-1719448

Abstract

Hand gesture recognition plays significant role in HCI applications and also it has been incorporated in the fields of Virtual Reality (VR), Augmented Reality (AR) and Robotics. Gestures in the real world are continuous and under these conditions, any recognition online would be very difficult since there are two types of gesture classes to be recognized as well as the temporal limits of these classes of gesture. Current state-of-the-art deep learning based on sliding windows and attention mechanisms yield piece-wise and “jagged” predictions resulting in more false alarms as well as incorrect gesture boundaries. Single-modality methods are not beneficial for recognition purposes because they do not return full-fledged gestures of the gesture. It is a reproduction of a cross-attention-based online hand gesture recognition model based on Joint Collection Distance (JCD) and Frame Vector (FV) that reproduces attention-based models. Taking the temporal refinement approach is proposed to reduce the noise and enhance the detection of the boundaries. To extend this into multimodal, visual appearance features of video frames are added in through an existing CNN which can learn structural and appearance features. Experimental findings of IPN Hand Gesture Dataset indicate that the Temporal refinement enhanced the Detection Rate of 0.9184 to 0.9288, False Positives of 39 to 14, and Mean IoU of 0.7375 to 0.8153 which showed that the gesture recognition and temporal localization were more accurate.

Keywords

Hand Gesture Recognition, Multimodal Learning, Attention Mechanism, Temporal Refinement, Joint Collection Distance (JCD), Frame Vector (FV), Deep Learning, Sliding Window.

How to cite this paper

Saiba Teja Sri, Sanaboina Chandra Sekhar "Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition" Iconic Research And Engineering Journals Volume 9 Issue 12 2026 Page 3661-3674 https://doi.org/10.64388/IREV9I12-1719448
Saiba Teja Sri, Sanaboina Chandra Sekhar "Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition" Iconic Research And Engineering Journals, vol. 9, no. 12, Jun. 2026, doi: https://doi.org/10.64388/IREV9I12-1719448
Saiba Teja Sri, Sanaboina Chandra Sekhar (2026). Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition. Iconic Research And Engineering Journals, 9(12). doi: https://doi.org/10.64388/IREV9I12-1719448
Saiba Teja Sri, Sanaboina Chandra Sekhar "Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition" Iconic Research And Engineering Journals, vol. 9, no. 12, Jun. 2026. Crossref, https://doi.org/10.64388/IREV9I12-1719448
@article{1719448,
      author = {Saiba Teja Sri, Sanaboina Chandra Sekhar},
      title = {Temporal And Multimodal Enhancement of Semantically Interpretable Attention for Online Hand Gesture Recognition},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {12},
      pages = {3661-3674},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1719448.pdf},
      abstract = {Hand gesture recognition plays significant role in HCI applications and also it has been incorporated in the fields of Virtual Reality (VR), Augmented Reality (AR) and Robotics. Gestures in the real world are continuous and under these conditions, any recognition online would be very difficult since there are two types of gesture classes to be recognized as well as the temporal limits of these classes of gesture. Current state-of-the-art deep learning based on sliding windows and attention mechanisms yield piece-wise and “jagged” predictions resulting in more false alarms as well as incorrect gesture boundaries. Single-modality methods are not beneficial for recognition purposes because they do not return full-fledged gestures of the gesture. It is a reproduction of a cross-attention-based online hand gesture recognition model based on Joint Collection Distance (JCD) and Frame Vector (FV) that reproduces attention-based models. Taking the temporal refinement approach is proposed to reduce the noise and enhance the detection of the boundaries. To extend this into multimodal, visual appearance features of video frames are added in through an existing CNN which can learn structural and appearance features. Experimental findings of IPN Hand Gesture Dataset indicate that the Temporal refinement enhanced the Detection Rate of 0.9184 to 0.9288, False Positives of 39 to 14, and Mean IoU of 0.7375 to 0.8153 which showed that the gesture recognition and temporal localization were more accurate.},
      keywords = {Hand Gesture Recognition, Multimodal Learning, Attention Mechanism, Temporal Refinement, Joint Collection Distance (JCD), Frame Vector (FV), Deep Learning, Sliding Window.},
      month = {June},
      doi = {https://doi.org/10.64388/IREV9I12-1719448}
  }