Identifying the Emotional State of Tourists in Rural Tourism Videos Using Multimodal Multi-Scale Attention Mechanisms
Main Article Content
Abstract
This study addresses the challenge of accurately identifying tourists’ emotional states from unstructured rural tourism videos under diverse environmental stimuli by incorporating multimodal analysis of visual, acoustic, and textual information. To achieve robust and fine-grained emotion recognition, a Multimodal Multi-Scale Attention Transformer (MMSAT) framework is proposed for collaborative modeling of audiovisual and verbal signals. The framework employs an intra-modal attention mechanism to adaptively fuse multi-scale visual representations while utilizing a three-layer Transformer encoder to capture deep cross-modal interactions and complementary semantic information. The fused representations are subsequently used to predict emotional valence and arousal with high precision. Experimental results on a self-constructed rural tourism video dataset demonstrate that the proposed MMSAT model achieves Concordance Correlation Coefficient (CCC) scores of 0.823 and 0.805 for valence and arousal prediction, respectively, outperforming the strongest baseline model, Multimodal Transformer (MulT), by 0.042 and 0.053. The proposed framework provides an effective solution for objective emotion assessment in complex open environments and offers methodological insights into multimodal information propagation and intelligent signal processing, with potential implications for advanced electromagnetic information perception and cross-domain sensing applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
L. Santamaria-Granados, J. F. Mendoza-Moreno, and G. Ramirez-Gonzalez, “Tourist recommender systems based on emotion recognition— a scientometric review,” Future Internet, vol. 13, no. 1, pp. 2, 2020, doi: 10.3390/fi13010002.
S. Zhao, G. Jia, J. Yang, G. Ding, and K. Keutzer, “Emotion recognition from multiple modalities: Fundamentals and methodologies,” IEEE Signal Processing Magazine, vol. 38, no. 6, pp. 59-73, 2021, doi: 10.1109/MSP.2021.3106895.
T. Zhang and Z. Tan, “Survey of deep emotion recognition in dynamic data using facial, speech and textual cues,” Multimedia Tools and Applications, vol. 83, no. 25, pp. 66223-66262, 2024, doi: 10.1007/s11042-023-17944-9.
H. Chen, M. Wang, and Z. Zhang, “Research on rural landscape preference based on TikTok short video content and user comments,” International Journal of Environmental Research and Public Health, vol. 19, no. 16, Art. no. 10115, 2022, doi: 10.3390/ijerph191610115.
S. K. Panda, A. K. Jena, M. R. Panda, and S. Panda, “Speech emotion recognition using multimodal feature fusion with machine learning approach,” Multimedia Tools and Applications, vol. 82, no. 27, pp. 42763-42781, 2023, doi: 10.1007/s11042-023-15275-3.
W. Nie, Y. Yan, D. Song, and K. Wang, “Multi-modal feature fusion based on multi-layers LSTM for video emotion recognition,” Multimedia Tools and Applications, vol. 80, no. 11, pp. 16205-16214, 2021, doi: 10.1007/s11042-020-08796-8.
M. Khan, P. N. Tran, N. T. Pham, A. El Saddik, and A. Othmani, “MemoCMT: Multimodal emotion recognition using cross-modal transformer-based feature fusion,” Scientific reports, vol. 15, no. 1, pp. 5473, 2025, doi: 10.1038/s41598-025-89202-x.
L. Zhao, Y. Yang, and T. Ning, “A Three-stage multimodal emotion recognition network based on text low-rank fusion,” Multimedia Systems, vol. 30, no. 3, pp. 142, 2024, doi: 10.1007/s00530-024-01345-5.
M. H. Yi, K. C. Kwak, and J. H. Shin, “HyFusER: Hybrid multimodal transformer for emotion recognition using dual cross modal attention,” Applied Sciences, vol. 15, no. 3, pp. 1053, 2025, doi: 10.3390/app15031053.
B. Xie, M. Sidulova, and C. H. Park, “Robust multimodal emotion recognition from conversation with transformerbased crossmodality fusion,” Sensors, vol. 21, no. 14, pp. 4913, 2021, doi: 10.3390/s21144913.
X. Zhang, M. Li, S. Lin, H. Xu, and G. Xiao, “Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 5, pp. 3192-3203, 2023, doi: 10.1109/TCSVT.2023.3312858.
A. Chaudhari, C. Bhatt, A. Krishna, and C. M. Travieso-González, “Facial emotion recognition with inter-modalityattention-transformer-based self-supervised learning,” Electronics, vol. 12, no. 2, pp. 288, 2023, doi: 10.3390/electronics12020288.
P. Bhattacharya, R. K. Gupta, and Y. Yang, “Exploring the contextual factors affecting multimodal emotion recognition in videos,” IEEE Transactions on Affective Computing, vol. 14, no. 2, pp. 1547-1557, 2021, doi: 10.1109/TAFFC.2021.3071503.
P. Yang, N. Liu, X. Liu, Y. Shu, W. Ji, Z. Ren, et al., “A multimodal dataset for mixed emotion recognition,” Scientific Data, vol. 11, no. 1, pp. 847, 2024, doi: 10.1038/s41597-024-03676-4.
W. Liu, J. L. Qiu, W. L. Zheng, and B. L. Lu, “Comparing recognition performance and robustness of multimodal deep learning models for multimodal emotion recognition,” IEEE Transactions on Cognitive and Developmental Systems, vol. 14, no. 2, pp. 715-729, 2021, doi: 10.1109/TCDS.2021.3071170.
C. Fang, H. Xiang, C. Leng, J. Chen, and Q. Yu, “Research on real-time detection of safety harness wearing of workshop personnel based on YOLOv5 and OpenPose,” Sustainability, vol. 14, no. 10, pp. 5872, 2022, doi: 10.3390/su14105872.
H. C. Nguyen, T. H. Nguyen, J. Nowak, A. Byrski, A. Siwocha, and V. Le, “Combined YOLOv5 and HRNet for high accuracy 2D keypoint and human pose estimation,” Journal of Artificial Intelligence and Soft Computing Research, vol. 12, no. 4, pp. 281-298, 2022, doi: 10.2478/jaiscr-2022-0019.
S. Mehra, V. Ranga, and R. Agarwal, “A deep learning approach to dysarthric utterance classification with BiLSTM-GRU, speech cue filtering, and log mel spectrograms,” The Journal of Supercomputing, vol. 80, no. 10, pp. 14520-14547, 2024, doi: 10.1007/s11227-024-06015-x.
G. Ahmed and A. A. Lawaye, “End-to-end ASR framework for Indian-English accent: Using speech CNN-based segmentation,” International Journal of Speech Technology, vol. 26, no. 4, pp. 903-918, 2023, doi: 10.1007/s10772-023-10053-w.
H. Boulal, M. Hamidi, M. Abarkan, and J. Barkani, “Amazigh CNN speech recognition system based on Mel spectrogram feature extraction method,” International Journal of Speech Technology, vol. 27, no. 1, pp. 287-296, 2024, doi: 10.1007/s10772-024-10100-0.
A. Dutta, G. Ashishkumar, and C. V. R. Rao, “Improving the Performance of ASR System by Building Acoustic Models using Spectro-Temporal and Phase-Based Features,” Circuits, Systems, and Signal Processing, vol. 41, no. 3, pp. 1609-1632, 2022, doi: 10.1007/s00034-021-01848-w.
X. Zhuang, F. Liu, J. Hou, J. Hao, and X. Cai, “Transformer-based interactive multi-modal attention network for video sentiment detection,” Neural Processing Letters, vol. 54, no. 3, pp. 1943-1960, 2022, doi: 10.1007/s11063-021-10713-5.
N. Pang, S. Guo, M. Yan, and C. A. Chan, “A short video classification framework based on cross-modal fusion,” Sensors, vol. 23, no. 20, pp. 8425, 2023, doi: 10.3390/s23208425.
L. H. Lee, J. H. Li, and L. C. Yu, “Chinese EmoBank: Building valence-arousal resources for dimensional sentiment analysis,” Transactions on Asian and Low-Resource Language Information Processing, vol. 21, no. 4, pp. 1-18, 2022, doi: 10.1145/3489141.
M. Yik, C. Mues, I. N. L. Sze, P. Kuppens, F. Tuerlinckx, K. De Roover, et al., “On the relationship between valence and arousal in samples across the globe,” Emotion, vol. 23, no. 2, pp. 332, 2023, doi: 10.1037/emo0001095.
G. S. Kumar, N. Sampathila, and R. J. Martis, “Classification of human emotional states based on valence-arousal scale using electroencephalogram,” Journal of Medical Signals & Sensors, vol. 13, no. 2, pp. 173-182, 2023, doi: 10.4103/jmss.jmss_169_21.
R. G. Praveen, P. Cardinal, and E. Granger, “Audio–visual fusion for emotion recognition in the valence–arousal space using joint cross-attention,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 5, no. 3, pp. 360-373, 2023, doi: 10.1109/TBIOM.2022.3233083.
M. Awatef, B. Hayet, and L. Zied, “Multimodal emotion recognition: Integrating speech and text for improved valence, arousal, and dominance prediction,” Annals of Telecommunications, vol. 80, no. 5, pp. 401-415, 2025, doi: 10.1007/s12243-025-01069-1.
J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, et al., “Dawn of the transformer era in speech emotion recognition: Closing the valence gap,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10745-10759, 2023, doi: 10.1109/TPAMI.2023.3263585.
C. Li, L. Xie, X. Wang, H. Pan, and Z. Wang, “A disentanglement mamba network with a temporally slack reconstruction mechanism for multimodal continuous emotion recognition,” Multimedia Systems, vol. 31, no. 2, pp. 1-18, 2025, doi: 10.1007/s00530-025-01758-w.
H. Lian, C. Lu, S. Li, Y. Zhao, C. Tang, and Y. Zong, “A survey of deep learning-based multimodal emotion recognition: Speech, text, and face,” Entropy, vol. 25, no. 10, pp. 1440, 2023, doi: 10.3390/e25101440.
X. Meng, “Cross-domain information fusion and personalized recommendation in artificial intelligence recommendation system based on mathematical matrix decomposition,” Scientific Reports, vol. 14, no. 1, pp. 7816, 2024, doi: 10.1038/s41598-024-57240-6.
W. Ai, Y. Shou, T. Meng, and K. Li, “Der-gcn: Dialog and event relation-aware graph convolutional neural network for multimodal dialog emotion recognition,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 3, pp. 4908-4921, 2024, doi: 10.1109/TNNLS.2024.3367940.
L. Guo, L. Wang, J. Dang, Y. Fu, J. Liu, and S. Ding, “Emotion recognition with multimodal transformer fusion framework based on acoustic and lexical information,” IEEE MultiMedia, vol. 29, no. 2, pp. 94-103, 2022, doi: 10.1109/MMUL.2022.3161411.
Y. Wu, M. Daoudi, and A. Amad, “Transformer-based self-supervised multimodal representation learning for wearable emotion recognition,” IEEE Transactions on Affective Computing, vol. 15, no. 1, pp. 157-172, 2023, doi: 10.1109/TAFFC.2023.3263907.