Home / Current Issue / Paper 1719374
Development of a Lightweight Temporal Convolutional Network for Microexpression Recognition
Subject area: Science,Engineering and Technology · Area of research: Deep Learning and Computer Vision
DOI: https://doi.org/10.64388/IREV10I1-1719374
Abstract
Microexpression recognition (MER) remains challenging because facial movements are short, subtle, and usually require computationally expensive temporal models. This paper presents a lightweight depthwise-separable Temporal Convolutional Network (DS-TCN) for MER on the CASME II dataset. The model uses a dual-stream design that combines grayscale facial-frame features with optical-flow motion features after face cropping, Eulerian video magnification, resizing, and SSIM-based key-frame selection. Standard temporal convolutions are replaced with depthwise separable temporal convolutions to reduce parameter count and floating-point operations while preserving sequence learning. Three-fold cross-validation produced a validation accuracy of 72.6% and a macro-F1 score of 67.58%. Compared with the standard TCN baseline, the proposed DS-TCN reduced parameters from 1.57 M to 0.219 M, FLOPs from 10.33 G to 1.50 G, and CPU inference time from 580 ms to 210 ms. The results show that lightweight temporal modeling can provide competitive MER performance with substantially lower computational cost.
Keywords
CASME II, Depthwise Separable Convolution, Microexpression Recognition, Optical Flow, Temporal Convolutional Network
References
[1] Aouayeb, M., Hamidouche, W., Soladié, C., Kpalma, K., & Séguier, R. (2021). Micro-expression recognition from local facial regions. Signal Processing: Image Communication, 99, 116457.
[2] Cai, X., Tang, H., & Chai, L. (2023). Micro expression recognition based on graph convolutional networks with LSTM. In 2023 35th Chinese Control and Decision Conference (CCDC) (pp. 5449-5453). IEEE.
[3] Gan, Y. S., Liong, S. T., Yau, W. C., Huang, Y. C., & Tan, L. K. (2019). OFF-ApexNet on micro-expression recognition system. Signal Processing: Image Communication, 74, 129-139.
[4] Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., & Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861.
[5] Jin, H., He, N., Li, Z., & Yang, P. (2024). Micro-expression recognition based on multi-scale 3D residual convolutional neural network. Mathematical Biosciences and Engineering, 21(4), 5007-5031.
[6] Khor, H. Q., See, J., Liong, S. T., Phan, R. C. W., & Lin, W. (2019). Dual-stream shallow networks for facial micro-expression recognition. In IEEE International Conference on Image Processing (pp. 36-40). IEEE.
[7] Kim, D. H., Baddar, W. J., & Ro, Y. M. (2016). Micro-expression recognition with expression-state constrained spatio-temporal feature representations. In ACM International Conference on Multimedia (pp. 382-386).
[8] Li, J., Wang, Y., See, J., & Liu, W. (2018). Micro-expression recognition based on 3D flow convolutional neural network. Pattern Analysis and Applications, 22, 1331-1339.
[9] Peng, W., Hong, X., Xu, Y., & Zhao, G. (2019). A boost in revealing subtle facial expressions: A consolidated Eulerian framework. In IEEE FG 2019 (pp. 1-5).
[10] Verma, M., Satish, M., Reddy, K., Meedimale, Y., & Kumar, S. (2021). AutoMER: Spatiotemporal neural architecture search for microexpression recognition. IEEE Transactions on Affective Computing.
[11] Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600-612.
[12] Wu, H.-Y., Rubinstein, M., Shih, E., Guttag, J., Durand, F., & Freeman, W. (2012). Eulerian video magnification for revealing subtle changes in the world. ACM Transactions on Graphics, 31(4), 1-8.
[13] Xue, F., Zhang, Y., Shao, Z., Cui, W., & Yuan, Z. (2025). HIM-PyraNet: Hierarchical attention and region-focused lightweight network for micro-expression recognition. Concurrency and Computation: Practice and Experience.
[14] Yu, Z., Chen, X., & Qu, C. (2024). SDGSA: A lightweight shallow dual-group symmetric attention network for micro-expression recognition. Complex & Intelligent Systems, 10, 8143-8162.
[15] Zhang, F., Huang, Z., Zhang, X., & Jin, Q. (2024). Adaptive temporal motion guided graph convolution network for micro-expression recognition. In IEEE International Conference on Multimedia and Expo.
[16] Zhang, L., & Arandjelović, O. (2021). Review of automatic microexpression recognition in the past decade. Machine Learning and Knowledge Extraction, 3(2), 414-434.
[17] Zong, Y., Zhang, M., & Zhao, G. (2020). Toward bridging microexpressions from different domains via domain adaptation methods. IEEE Transactions on Affective Computing, 13, 1117-1127.
How to cite this paper
@article{1719374,
author = {Adeyemi, Ifeoluwa Olamiji, Falohun, Adeleye Samuel, Oguntoye Jonathan Ponmile, Ayinla, Michael Oluwaseun},
title = {Development of a Lightweight Temporal Convolutional Network for Microexpression Recognition},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {1},
pages = {1644-1655},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1719374.pdf},
abstract = {Microexpression recognition (MER) remains challenging because facial movements are short, subtle, and usually require computationally expensive temporal models. This paper presents a lightweight depthwise-separable Temporal Convolutional Network (DS-TCN) for MER on the CASME II dataset. The model uses a dual-stream design that combines grayscale facial-frame features with optical-flow motion features after face cropping, Eulerian video magnification, resizing, and SSIM-based key-frame selection. Standard temporal convolutions are replaced with depthwise separable temporal convolutions to reduce parameter count and floating-point operations while preserving sequence learning. Three-fold cross-validation produced a validation accuracy of 72.6% and a macro-F1 score of 67.58%. Compared with the standard TCN baseline, the proposed DS-TCN reduced parameters from 1.57 M to 0.219 M, FLOPs from 10.33 G to 1.50 G, and CPU inference time from 580 ms to 210 ms. The results show that lightweight temporal modeling can provide competitive MER performance with substantially lower computational cost.},
keywords = {CASME II, Depthwise Separable Convolution, Microexpression Recognition, Optical Flow, Temporal Convolutional Network},
month = {July},
doi = {https://doi.org/10.64388/IREV10I1-1719374}
}