Spatial Transformer Networks for Facial Action Unit Detection: Modeling Inter-AU Dependencies via Structured Attention

Main Article Content

X. Y. Yang
K. Wang

Abstract

Facial Action Unit (AU) detection is fundamental to understanding human facial expressions and emotions. While Transformers have revolutionized computer vision through self-attention mechanisms, their application to AU detection has primarily followed sequential processing paradigms. In this paper, we present a novel Spatial Transformer framework that reimagines AU relationship modeling through structured spatial attention mechanisms. Unlike conventional Vision Transformers that process images as sequences of patches, our approach treats AUs as spatially and functionally related entities, leveraging anatomically-informed attention patterns to capture complex inter-AU dependencies. The key innovation lies in our structured attention mechanism that incorporates multi-dimensional relationship encodings, enabling the model to learn both co-occurrence patterns and mutual exclusion constraints among AUs. We introduce spatial position encodings based on facial muscle locations rather than learnable embeddings, providing stronger inductive biases for facial analysis tasks. Our comprehensive experiments demonstrate that the Spatial Transformer effectively captures subtle AU relationships that are often missed by independent detection approaches. The learned attention patterns align remarkably well with established facial anatomy knowledge and FACS-defined AU relationships, offering interpretable insights into the model’s decision-making process. Furthermore, our analysis reveals that spatial attention mechanisms are particularly effective for detecting AUs with low activation intensities and those that frequently co-occur with other units. This work not only advances the state-of-the-art in AU detection but also provides a new perspective on modeling structured relationships in facial analysis through the lens of spatial attention, opening avenues for more sophisticated facial behavior understanding systems.

Downloads

Download data is not yet available.

Article Details

How to Cite
Yang, X. Y., & Wang, K. (2026). Spatial Transformer Networks for Facial Action Unit Detection: Modeling Inter-AU Dependencies via Structured Attention. Advanced Electromagnetics, 15(3), 9761–9771. https://doi.org/10.7716/aem.v15i3.4169
Section
Research Articles

References

A. Vaswani et al., “Attention is all you need,” Advances in neural information processing systems, 30, 2017.

J. Devlin, M.-W. Chang, K. Lee and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, vol. 1 (long and short papers), pp. 4171–4186, 2019.

T. Brown et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020.

A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020.

Z. Liu, “Swin transformer: Hierarchical vision transformer using shifted windows,” in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 10012–10022, 2021.

H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles and H. J’egou, “Training data-efficient image transformers & distillation through attention,” in: International conference on machine learning, PMLR, pp. 10347–10357, 2021.

C.-F. R. Chen, Q. Fan and R. Panda, “Crossvit: Crossattention multi-scale vision transformer for image classification,” in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 357–366, 2021.

B. Martinez, M. F. Valstar, B. Jiang and M. Pantic, “Automatic analysis of facial actions: A survey,” IEEE transactions on affective computing, vol. 10, no. 3, pp. 325–347, 2017.

P. Ekman and W. V. Friesen, “Facial action coding system,” Environmental Psychology & Nonverbal Behavior, 1978.

T. Baltruˇsaitis, P. Robinson and L.-P. Morency, “Openface: an open source facial behavior analysis toolkit,” in: 2016 IEEE winter conference on applications of computer vision (WACV), IEEE, pp. 1–10, 2016.

J. Xia, W. Qu, W. Huang, J. Zhang, X. Wang and M. Xu, Sparse local patch transformer for robust face alignment and landmarks inherent relation learning,” in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4052–4061, 2022.

X. Chu, Z. Tian, B. Zhang, X. Wang and C. Shen, Conditional positional encodings for vision transformers,” arXiv preprint arXiv:2102.10882, 2021.

C.-F. Chen, R. Panda and Q. Fan, “Regionvit: Regionalto-local attention for vision transformers,” arXiv preprint arXiv:2106.02689, 2021.

G. M. Jacob, B. Stenger, “Facial action unit detection with transformers,” in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 7680–7689, 2021.

K. Yuan, Z. Yu, X. Liu, W. Xie, H. Yue and J. Yang, Auformer: Vision transformers are parameterefficient facial action unit detectors,” in: European Conference on Computer Vision, Springer, pp. 427–445, 2024.

Z. Shao, Z. Liu, J. Cai and L. Ma, “Jaa-net: Joint facial action unit detection and face alignment via adaptive attention,” International Journal of Computer Vision, vol. 129, pp. 321–340, 2021.

X. Zhang et al., “Bp4dspontaneous: a high-resolution spontaneous 3d dynamic facial expression database,” Image and Vision Computing, vol. 32, no. 10, pp. 692–706, 2014.

S. M. Mavadati, M. H. Mahoor, K. Bartlett, P. Trinh and J. F. Cohn, “Disfa: A spontaneous facial action intensity database,” IEEE Transactions on Affective Computing, vol. 4, no. 2, pp. 151–160, 2013.

G.-B. Duchenne, “The mechanism of human facial expression,” Cambridge university press, 1990.

Y.-I. Tian, T. Kanade and J. F. Cohn, “Recognizing action units for facial expression analysis,” IEEE Transactions on pattern analysis and machine intelligence, vol. 23, no. 2 pp. 97–115, 2001.

T. Ahonen, A. Hadid and M. Pietik¨ainen, “Face recognition with local binary patterns,” in: Computer Vision-ECCV 2004: 8th European Conference on Computer Vision, Prague, Czech Republic, May 11-14, 2004. Proceedings, Part I 8, Springer, pp. 469–481, 2004.

Y. Bai, L. Guo, L. Jin and Q. Huang, A novel feature extraction method using pyramid histogram of orientation gradients for smile recognition, in: 2009 16th IEEE International Conference on Image Processing (ICIP), IEEE, pp. 3305– 3308, 2009.

B. Jiang, M. F. Valstar and M. Pantic, “Action unit detection using sparse appearance descriptors in spacetime video volumes,” in: 2011 IEEE International Conference on Automatic Face & Gesture Recognition (FG), IEEE, pp. 314– 321, 2011.

K. Zhao, W.-S. Chu and H. Zhang, Deep region and multi-label learning for facial action unit detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3391–3399, 2016.

W. Li, F. Abtahi, Z. Zhu and L. Yin, “Eac-net: A region-based deep enhancing and cropping approach for facial action unit detection,” in: 2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017), IEEE, pp. 103–110, 2017.

W. Li, F. Abtahi, Z. Zhu and L. Yin, “Eac-net: Deep nets with enhancing and cropping for facial action unit detection,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 11, pp. 2583–2596, 2018.

C. Corneanu, M. Madadi and S. Escalera, “Deep structure inference network for facial action unit recognition,” in: Proceedings of the european conference on computer vision (ECCV), pp. 298–313, 2018.

Z. Shao, Z. Liu, J. Cai and L. Ma, “Deep adaptive attention for joint facial action unit detection and face alignment,” in: Proceedings of the European conference on computer vision (ECCV), pp. 705–720, 2018.

X. Niu, H. Han, S. Yang, Y. Huang and S. Shan, “Local relationship learning with person-specific shape regularization for facial action unit detection,” in: Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pp. 11917–11926, 2019.

Z. Shao, Z. Liu, J. Cai, Y. Wu and L. Ma, “Facial action unit detection using attention and relation learning,” IEEE transactions on affective computing, vol. 13, no. 3, pp. 1274–1289, 2019.

H. Yang, L. Yin, Y. Zhou and J. Gu, “Exploiting semantic embedding and visual feature for facial action unit detection,” in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10482–10491, 2021.

Y. Chen, D. Chen, T. Wang, Y. Wang and Y. Liang, “Causal intervention for subject-deconfounded facial action unit recognition,” in: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 374–382, 2022.

C. Bisogni, A. Castiglione, S. Hossain, F. Narducci and S. Umer, “Impact of deep learning approaches on facial expression recognition in healthcare industries,” IEEE Transactions on Industrial Informatics, vol. 18, no. 8, pp. 5619–5627, 2022.

R. Daza et al., “Matt: Multimodal attention level estimation for e-learning platforms,” arXiv preprint arXiv:2301.09174, 2023.

V. G. Prakash et al., “Computer vision-based assessment of autistic children: Analyzing interactions, emotions, human pose, and life skills,” IEEE Access, vol. 11, pp. 47907–47929, 2023.

M. Ahmad, et al., “Facial expression recognition using lightweight deep learning modeling,” Mathematical Biosciences and Engineering, vol. 20, no. 5, pp. 8208, 2023.

A. Mehrabian, et al., Silent messages, Wadsworth Belmont, CA, 1971.

G. Li, X. Zhu, Y. Zeng, Q. Wang and L. Lin, “Semantic relationships guided representation learning for facial action unit recognition,” in: Proceedings of the AAAI conference on artificial intelligence, vol. 33, pp. 8594–8601, 2019.

T. Song, L. Chen, W. Zheng and Q. Ji, “Uncertain graph neural networks for facial action unit detection,” in: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 5993–6001, 2021.

T. Song, Z. Cui, W. Zheng and Q. Ji, “Hybrid message passing with performance-driven structures for facial action unit detection,” in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6267–6276, 2021.

H. Wang, B. Li, S. Wu, S. Shen, F. Liu and S. Ding, A. Zhou, Rethinking the learning paradigm for dynamic facial expression recognition,” in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 17958–17968, 2023.

P. Melzi, et al., “Frcsyn challenge at wacv 2024: Face recognition challenge in the era of synthetic data,” in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 892–901, 2024.

K. He, X. Zhang, S. Ren and J. Sun, “Deep residual learning for image recognition,” in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.