Identifying the Emotional State of Tourists in Rural Tourism Videos Using Multimodal Multi-Scale Attention Mechanisms

Main Article Content

Z. G. Yang

Abstract

This study addresses the challenge of accurately identifying tourists’ emotional states from unstructured rural tourism videos under diverse environmental stimuli by incorporating multimodal analysis of visual, acoustic, and textual information. To achieve robust and fine-grained emotion recognition, a Multimodal Multi-Scale Attention Transformer (MMSAT) framework is proposed for collaborative modeling of audiovisual and verbal signals. The framework employs an intra-modal attention mechanism to adaptively fuse multi-scale visual representations while utilizing a three-layer Transformer encoder to capture deep cross-modal interactions and complementary semantic information. The fused representations are subsequently used to predict emotional valence and arousal with high precision. Experimental results on a self-constructed rural tourism video dataset demonstrate that the proposed MMSAT model achieves Concordance Correlation Coefficient (CCC) scores of 0.823 and 0.805 for valence and arousal prediction, respectively, outperforming the strongest baseline model, Multimodal Transformer (MulT), by 0.042 and 0.053. The proposed framework provides an effective solution for objective emotion assessment in complex open environments and offers methodological insights into multimodal information propagation and intelligent signal processing, with potential implications for advanced electromagnetic information perception and cross-domain sensing applications.

Downloads

Download data is not yet available.

Article Details

How to Cite
Yang, Z. G. (2026). Identifying the Emotional State of Tourists in Rural Tourism Videos Using Multimodal Multi-Scale Attention Mechanisms. Advanced Electromagnetics, 15(3), 5638–5650. https://doi.org/10.7716/aem.v15i3.3615
Section
Research Articles

References

L. Santamaria-Granados, J. F. Mendoza-Moreno, and G. Ramirez-Gonzalez, “Tourist recommender systems based on emotion recognition— a scientometric review,” Future Internet, vol. 13, no. 1, pp. 2, 2020, doi: 10.3390/fi13010002.

View Article

S. Zhao, G. Jia, J. Yang, G. Ding, and K. Keutzer, “Emotion recognition from multiple modalities: Fundamentals and methodologies,” IEEE Signal Processing Magazine, vol. 38, no. 6, pp. 59-73, 2021, doi: 10.1109/MSP.2021.3106895.

View Article

T. Zhang and Z. Tan, “Survey of deep emotion recognition in dynamic data using facial, speech and textual cues,” Multimedia Tools and Applications, vol. 83, no. 25, pp. 66223-66262, 2024, doi: 10.1007/s11042-023-17944-9.

View Article

H. Chen, M. Wang, and Z. Zhang, “Research on rural landscape preference based on TikTok short video content and user comments,” International Journal of Environmental Research and Public Health, vol. 19, no. 16, Art. no. 10115, 2022, doi: 10.3390/ijerph191610115.

View Article

S. K. Panda, A. K. Jena, M. R. Panda, and S. Panda, “Speech emotion recognition using multimodal feature fusion with machine learning approach,” Multimedia Tools and Applications, vol. 82, no. 27, pp. 42763-42781, 2023, doi: 10.1007/s11042-023-15275-3.

View Article

W. Nie, Y. Yan, D. Song, and K. Wang, “Multi-modal feature fusion based on multi-layers LSTM for video emotion recognition,” Multimedia Tools and Applications, vol. 80, no. 11, pp. 16205-16214, 2021, doi: 10.1007/s11042-020-08796-8.

View Article

M. Khan, P. N. Tran, N. T. Pham, A. El Saddik, and A. Othmani, “MemoCMT: Multimodal emotion recognition using cross-modal transformer-based feature fusion,” Scientific reports, vol. 15, no. 1, pp. 5473, 2025, doi: 10.1038/s41598-025-89202-x.

View Article

L. Zhao, Y. Yang, and T. Ning, “A Three-stage multimodal emotion recognition network based on text low-rank fusion,” Multimedia Systems, vol. 30, no. 3, pp. 142, 2024, doi: 10.1007/s00530-024-01345-5.

View Article

M. H. Yi, K. C. Kwak, and J. H. Shin, “HyFusER: Hybrid multimodal transformer for emotion recognition using dual cross modal attention,” Applied Sciences, vol. 15, no. 3, pp. 1053, 2025, doi: 10.3390/app15031053.

View Article

B. Xie, M. Sidulova, and C. H. Park, “Robust multimodal emotion recognition from conversation with transformerbased crossmodality fusion,” Sensors, vol. 21, no. 14, pp. 4913, 2021, doi: 10.3390/s21144913.

View Article

X. Zhang, M. Li, S. Lin, H. Xu, and G. Xiao, “Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 5, pp. 3192-3203, 2023, doi: 10.1109/TCSVT.2023.3312858.

View Article

A. Chaudhari, C. Bhatt, A. Krishna, and C. M. Travieso-González, “Facial emotion recognition with inter-modalityattention-transformer-based self-supervised learning,” Electronics, vol. 12, no. 2, pp. 288, 2023, doi: 10.3390/electronics12020288.

View Article

P. Bhattacharya, R. K. Gupta, and Y. Yang, “Exploring the contextual factors affecting multimodal emotion recognition in videos,” IEEE Transactions on Affective Computing, vol. 14, no. 2, pp. 1547-1557, 2021, doi: 10.1109/TAFFC.2021.3071503.

View Article

P. Yang, N. Liu, X. Liu, Y. Shu, W. Ji, Z. Ren, et al., “A multimodal dataset for mixed emotion recognition,” Scientific Data, vol. 11, no. 1, pp. 847, 2024, doi: 10.1038/s41597-024-03676-4.

View Article

W. Liu, J. L. Qiu, W. L. Zheng, and B. L. Lu, “Comparing recognition performance and robustness of multimodal deep learning models for multimodal emotion recognition,” IEEE Transactions on Cognitive and Developmental Systems, vol. 14, no. 2, pp. 715-729, 2021, doi: 10.1109/TCDS.2021.3071170.

View Article

C. Fang, H. Xiang, C. Leng, J. Chen, and Q. Yu, “Research on real-time detection of safety harness wearing of workshop personnel based on YOLOv5 and OpenPose,” Sustainability, vol. 14, no. 10, pp. 5872, 2022, doi: 10.3390/su14105872.

View Article

H. C. Nguyen, T. H. Nguyen, J. Nowak, A. Byrski, A. Siwocha, and V. Le, “Combined YOLOv5 and HRNet for high accuracy 2D keypoint and human pose estimation,” Journal of Artificial Intelligence and Soft Computing Research, vol. 12, no. 4, pp. 281-298, 2022, doi: 10.2478/jaiscr-2022-0019.

View Article

S. Mehra, V. Ranga, and R. Agarwal, “A deep learning approach to dysarthric utterance classification with BiLSTM-GRU, speech cue filtering, and log mel spectrograms,” The Journal of Supercomputing, vol. 80, no. 10, pp. 14520-14547, 2024, doi: 10.1007/s11227-024-06015-x.

View Article

G. Ahmed and A. A. Lawaye, “End-to-end ASR framework for Indian-English accent: Using speech CNN-based segmentation,” International Journal of Speech Technology, vol. 26, no. 4, pp. 903-918, 2023, doi: 10.1007/s10772-023-10053-w.

View Article

H. Boulal, M. Hamidi, M. Abarkan, and J. Barkani, “Amazigh CNN speech recognition system based on Mel spectrogram feature extraction method,” International Journal of Speech Technology, vol. 27, no. 1, pp. 287-296, 2024, doi: 10.1007/s10772-024-10100-0.

View Article

A. Dutta, G. Ashishkumar, and C. V. R. Rao, “Improving the Performance of ASR System by Building Acoustic Models using Spectro-Temporal and Phase-Based Features,” Circuits, Systems, and Signal Processing, vol. 41, no. 3, pp. 1609-1632, 2022, doi: 10.1007/s00034-021-01848-w.

View Article

X. Zhuang, F. Liu, J. Hou, J. Hao, and X. Cai, “Transformer-based interactive multi-modal attention network for video sentiment detection,” Neural Processing Letters, vol. 54, no. 3, pp. 1943-1960, 2022, doi: 10.1007/s11063-021-10713-5.

View Article

N. Pang, S. Guo, M. Yan, and C. A. Chan, “A short video classification framework based on cross-modal fusion,” Sensors, vol. 23, no. 20, pp. 8425, 2023, doi: 10.3390/s23208425.

View Article

L. H. Lee, J. H. Li, and L. C. Yu, “Chinese EmoBank: Building valence-arousal resources for dimensional sentiment analysis,” Transactions on Asian and Low-Resource Language Information Processing, vol. 21, no. 4, pp. 1-18, 2022, doi: 10.1145/3489141.

View Article

M. Yik, C. Mues, I. N. L. Sze, P. Kuppens, F. Tuerlinckx, K. De Roover, et al., “On the relationship between valence and arousal in samples across the globe,” Emotion, vol. 23, no. 2, pp. 332, 2023, doi: 10.1037/emo0001095.

View Article

G. S. Kumar, N. Sampathila, and R. J. Martis, “Classification of human emotional states based on valence-arousal scale using electroencephalogram,” Journal of Medical Signals & Sensors, vol. 13, no. 2, pp. 173-182, 2023, doi: 10.4103/jmss.jmss_169_21.

View Article

R. G. Praveen, P. Cardinal, and E. Granger, “Audio–visual fusion for emotion recognition in the valence–arousal space using joint cross-attention,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 5, no. 3, pp. 360-373, 2023, doi: 10.1109/TBIOM.2022.3233083.

View Article

M. Awatef, B. Hayet, and L. Zied, “Multimodal emotion recognition: Integrating speech and text for improved valence, arousal, and dominance prediction,” Annals of Telecommunications, vol. 80, no. 5, pp. 401-415, 2025, doi: 10.1007/s12243-025-01069-1.

View Article

J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, et al., “Dawn of the transformer era in speech emotion recognition: Closing the valence gap,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10745-10759, 2023, doi: 10.1109/TPAMI.2023.3263585.

View Article

C. Li, L. Xie, X. Wang, H. Pan, and Z. Wang, “A disentanglement mamba network with a temporally slack reconstruction mechanism for multimodal continuous emotion recognition,” Multimedia Systems, vol. 31, no. 2, pp. 1-18, 2025, doi: 10.1007/s00530-025-01758-w.

View Article

H. Lian, C. Lu, S. Li, Y. Zhao, C. Tang, and Y. Zong, “A survey of deep learning-based multimodal emotion recognition: Speech, text, and face,” Entropy, vol. 25, no. 10, pp. 1440, 2023, doi: 10.3390/e25101440.

View Article

X. Meng, “Cross-domain information fusion and personalized recommendation in artificial intelligence recommendation system based on mathematical matrix decomposition,” Scientific Reports, vol. 14, no. 1, pp. 7816, 2024, doi: 10.1038/s41598-024-57240-6.

View Article

W. Ai, Y. Shou, T. Meng, and K. Li, “Der-gcn: Dialog and event relation-aware graph convolutional neural network for multimodal dialog emotion recognition,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 3, pp. 4908-4921, 2024, doi: 10.1109/TNNLS.2024.3367940.

View Article

L. Guo, L. Wang, J. Dang, Y. Fu, J. Liu, and S. Ding, “Emotion recognition with multimodal transformer fusion framework based on acoustic and lexical information,” IEEE MultiMedia, vol. 29, no. 2, pp. 94-103, 2022, doi: 10.1109/MMUL.2022.3161411.

View Article

Y. Wu, M. Daoudi, and A. Amad, “Transformer-based self-supervised multimodal representation learning for wearable emotion recognition,” IEEE Transactions on Affective Computing, vol. 15, no. 1, pp. 157-172, 2023, doi: 10.1109/TAFFC.2023.3263907.

View Article

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > >> 

You may also start an advanced similarity search for this article.