Exploring the Differences in Audience Emotional Arousal Patterns Between Virtual and Real-Life Anchors Using BERT Multimodal Sentiment Analysis
Main Article Content
Abstract
This study investigates audience emotional arousal patterns in virtual versus real-life anchors using a multi-channel signal processing framework. A Text-Anchored Cross-modal Alignment for Arousal Modeling (TACAM) method is proposed, integrating text, visual, and audio signals from live streaming platforms. Textual features are extracted via fine-tuned BERT-large, while visual motion energy and audio prosodic features are temporally aligned and fused through a cross-modal attention mechanism. A Transformer-based temporal regression network generates high-resolution emotional arousal curves, capturing dynamic intensity, rise rate, and high-arousal duration. Experimental results show that the model achieves a high-arousal duration ratio of 0.38 for virtual anchors and 0.52 for real-life anchors, with adaptive weighting reflecting modality reliability (text weight 76% for virtual anchors). By interpreting audience reactions as multi-channel signal flows and TACAM as a signal fusion and temporal alignment system with closed-loop modulation, this work provides an engineering-oriented perspective for designing robust, interpretable, and quantitative frameworks for human-computer emotional interaction, enabling more accurate modeling of real-time affective dynamics in digital media environments.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
J. Ham, S. Li, J. Looi, and M. S. Eastin, “Virtual humans as social actors: Investigating user perceptions of virtual humans’ emotional expression on social media,” Computers in Human Behavior, vol. 155, Art. no. 108161, 2024, doi: 10.1016/j.chb.2024.108161.
X. Liu and L. Zhang, “Virtual streamer responsiveness and consumer purchase intention in live streaming commerce: Social presence as a mediator,” Social Behavior and Personality: An international journal, vol. 52, no. 11, pp. 1-7, 2024, doi: 10.2224/sbp.13708.
Y. Yuan, S. Duo, X. Tong, and Y. Wang, “A multimodal affective interaction architecture integrating BERT-based semantic understanding and VITS-based emotional speech synthesis,” Algorithms, vol. 18, no. 8, pp. 513, 2025, doi: 10.3390/a18080513.
R. Das and T. D. Singh, “Multimodal sentiment analysis: A survey of methods, trends, and challenges,” ACM Computing Surveys, vol. 55, no. 13s, pp. 1-38, 2023, doi: 10.1145/3586075.
R. Li, X. Yang, J. Lou, and J. Zhang, “A temporal-spectral graph convolutional neural network model for EEG emotion recognition within and across subjects,” Brain Informatics, vol. 11, no. 1, pp. 30, 2024, doi: 10.1186/s40708-024-00242-x.
C. Lee, H. Lee, and M. Whang, “Emotion Recognition from rPPG via Physiologically Inspired Temporal Encoding and Attention-Based Curriculum Learning,” Sensors, vol. 25, no. 13, pp. 3995, 2025, doi: 10.3390/s25133995.
Z. Zhang, W. Wu, T. Yuan, and G. Feng, “Modality-Enhanced Multimodal Integrated Fusion Attention Model for Sentiment Analysis,” Applied Sciences, vol. 15, no. 19, Art. no. 10825, 2025, doi: 10.3390/app151910825.
Z. Li, P. Liu, Y. Pan, J. Yu, W. Liu, H. Chen, et al., “Text-dominant multimodal perception network for sentiment analysis based on cross-modal semantic enhancements,” Applied Intelligence, vol. 55, pp. 188, 2025, doi: 10.1007/s10489-024-06150-1.
Y. Song, L. Zhang, R. Zhang, H. Zhan, M. Dai, X. Hu, et al., “MSF-Net: A Data-Driven Multimodal Transformer for Intelligent Behavior Recognition and Financial Risk Reasoning in Virtual Live-Streaming,” Electronics, vol. 14, no. 23, pp. 4769, 2025, doi: 10.3390/electronics14234769.
R. Abbas, B. W. Schuller, X. Li, C. Lin, and X. Li, “Emotion recognition in live broadcasting: A multimodal deep learning framework,” Multimedia Systems, vol. 31, pp. 253, 2025, doi: 10.1007/s00530-025-01780-y.
U. K. Das, R. S. Ani, N. Datta, I. Fahad, J. Sikder, U. Sara, et al., “Enhancing sentiment analysis accuracy on social media comments using a tuned BERT model,” Discover Computing, vol. 28, no. 1, pp. 198, 2025, doi: 10.1007/s10791-025-09599-x.
H. Peng, X. Gu, J. Li, Z. Wang, and H. Xu, “Text-centric multimodal contrastive learning for sentiment analysis,” Electronics, vol. 13, no. 6, pp. 1149, 2024, doi: 10.3390/electronics13061149.
S. Bai, L. Cao, and J. Zhou, “Is the Anthropomorphic Virtual Anchor Its Optimal Form? An Exploration of the Impact of Virtual Anchors’ Appearance on Consumers’ Emotions and Purchase Intention,” Journal of Theoretical and Applied Electronic Commerce Research, vol. 20, no. 2, pp. 110, 2025, doi: 10.3390/jtaer20020110.
D. Wang, Y. Peng, L. Haddouk, N. Vayatis, and P. P. Vidal, “Assessing virtual reality presence through physiological measures: A comprehensive review,” Frontiers in Virtual Reality, vol. 6, Art. no. 1530770, 2025, doi: 10.3389/frvir.2025.1530770.
O. Galal, A. H. Abdel-Gawad, and M. Farouk, “Rethinking of BERT sentence embedding for text classification,” Neural Computing and Applications, vol. 36, pp. 20245-20258, 2024, doi: 10.1007/s00521-024-10212-3.
A. S. Talaat, “Sentiment analysis classification system using hybrid BERT models,” Journal of Big Data, vol. 10, no. 1, pp. 110, 2023, doi: 10.1186/s40537-023-00781-w.
X. Wang, H. Ren, and A. Wang, “Smish: A novel activation function for deep learning methods,” Electronics, vol. 11, no. 4, pp. 540, 2022, doi: 10.3390/electronics11040540.
Y. R. Youn and J. K. Hong, “Enhancing convolutional neural network performance through optimized sigmoid activation function modeling,” Asia-Pacific Journal of Convergent Research Interchange, vol. 9, no. 10, pp. 51-59, 2023, doi: 10.47116/apjcri.2023.10.05.
A. C. Cob-Parro, C. Losada-Gutierrez, and M. Marron-Romera, “Stampede detector based on deep learning models using dense optical flow,” Engineering Applications of Artificial Intelligence, vol. 142, Art. no. 109940, 2025, doi: 10.1016/j.engappai.2024.109940.
N. Ranjan, S. Bhandari, Y. Kim, and H. Kim, “Video Frame Prediction by Joint Optimization of Direct Frame Synthesis and Optical-Flow Estimation,” Computers, Materials and Continua, vol. 75, no. 2, pp. 2615-2639, 2023, doi: 10.32604/cmc.2023.026086.
M. G. Huddar, S. S. Sannakki, and V. S. Rajpurohit, “Attention-based multimodal contextual fusion for sentiment and emotion classification using bidirectional LSTM,” Multimedia Tools and Applications, vol. 80, no. 9, pp. 13059-13076, 2021, doi: 10.1007/s11042-020-10285-x.
X. Zhuang, F. Liu, J. Hou, J. Hao, and X. Cai, “Transformer-based interactive multi-modal attention network for video sentiment detection,” Neural Processing Letters, vol. 54, pp. 1943-1960, 2022, doi: 10.1007/s11063-021-10713-5.
C. Shi and Y. Zhang, “MMKT: Multimodal Sentiment Analysis Model Based on Knowledge-Enhanced and Text-Guided Learning,” Applied Sciences, vol. 15, no. 17, pp. 9815, 2025, doi: 10.3390/app15179815.
J. Hou, N. Omar, S. Tiun, S. Saad, and Q. He, “TCHFN: Multimodal sentiment analysis based on text-centric hierarchical fusion network,” Knowledge-Based Systems, vol. 300, no. 1, Art. no. 112220, 2024, doi: 10.1016/j.knosys.2024.112220.
C. Liu, Y. Wang, and J. Yang, “A transformer-encoder-based multimodal multi-attention fusion network for sentiment analysis,” Applied Intelligence, vol. 54, no. 17-18, pp. 8415-8441, 2024, doi: 10.1007/s10489-024-05623-7.
N. M. Foumani, C. W. Tan, G. I. Webb, and M. Salehi, “Improving position encoding of transformers for multivariate time series classification,” Data Mining and Knowledge Discovery, vol. 38, no. 1, pp. 22-48, 2024, doi: 10.1007/s10618-023-00948-2.
Y. Dogan, “A new global pooling method for deep neural networks: Global average of top-k max-pooling,” Traitement du Signal, vol. 40, no. 2, pp. 577-587, 2023, doi: 10.18280/ts.400216.
A. Zafar, M. Aamir, N. Mohd Nawi, A. Arshad, S. Riaz, A. Alruban, et al., “A comparison of pooling methods for convolutional neural networks,” Applied Sciences, vol. 12, no. 17, pp. 8643, 2022, doi: 10.3390/app12178643.
K. Zhao, M. Zheng, Q. Li, and J. Liu, “Multimodal Sentiment Analysis— A Comprehensive Survey From a Fusion Methods Perspective,” IEEE Access, vol. 13, no. 1, pp. 64556-64583, 2025, doi: 10.1109/ACCESS.2025.3554665.
X. Yan, H. Xue, S. Jiang, and Z. Liu, “Multimodal sentiment analysis using multi-tensor fusion network with crossmodal modeling,” Applied Artificial Intelligence, vol. 36, no. 1, Art. no. 2000688, 2022, doi: 10.1080/08839514.2021.2000688.
D. Wang, S. Liu, Q. Wang, Y. Tian, L. He, and X. Gao, “Cross-modal enhancement network for multimodal sentiment analysis,” IEEE Transactions on Multimedia, vol. 25, pp. 4909-4921, 2023, doi: 10.1109/TMM.2022.3183830.