Classification of Highly-Rated Comments in Rural Short Videos Based on Multimodal Emotion Features
Main Article Content
Abstract
The rapid growth of rural short video platforms has generated large volumes of multimodal interaction data, where audiovisual signals and textual comments jointly reflect users’ emotional preferences and engagement patterns. However, semantic inconsistency between heterogeneous modalities and the lack of interpretable emotion modeling limit the accurate classification of highly-rated comments. This study proposes a cross-modal emotion alignment framework that integrates temporal audiovisual features with textual semantic representations for fine-grained comment classification. A dual-stream encoding network combines temporal attention and Wav2Vec2.0 to capture scene evolution, dialect rhythms, and ambient acoustic characteristics, while a domain-adapted text encoder extracts implicit emotional semantics from user comments. A bidirectional cross-attention gating module dynamically aligns multimodal features and suppresses semantic noise caused by irony and metaphor, followed by a multi-head classifier for structured emotion prediction. Experiments on a self-constructed rural short video corpus demonstrate that the proposed model achieves a Macro-F1 score of 89.4% while reducing the confusion rate to 11.9%, outperforming representative baseline methods. Beyond improving multimodal sentiment analysis, the proposed architecture establishes an effective signal-level fusion and feature interaction strategy for heterogeneous audiovisual data, providing useful insights for intelligent information perception, multimodal signal transmission, and adaptive communication systems in electromagnetic sensing and next-generation interactive media applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
J. Yu, “A Study on the narrative space of rural short video: A case study of rural short video in TikTok platform,” Frontiers in Art Research, vol. 5, no. 2, pp. 6-10, 2023, doi: 10.25236/FAR.2023.050202.
S. Zou, “Curating a scopic contact zone: Short video, rural performativity, and the mediatization of socio-spatial order in China,” Television & New Media, vol. 24, no. 4, pp. 452-470, 2023, doi: 10.1177/15274764221128925.
Z. Shuang, “Rural Short Videos’ New Farmer Image Construction and Communication Strategy Presentation–the Example of“ Winning the Heart,” Academic Journal of Humanities & Social Sciences, vol. 6, no. 14, pp. 63-68, 2023, doi: 10.25236/AJHSS.2023.061411.
D. Huo, P. Zou, and Y. Lu, “Live vs,” static comments: Empirical analysis of their differential effects on user evaluation of online videos. Journal of Theoretical and Applied Electronic Commerce Research, vol. 20, no. 2, pp. 102-125, 2025, doi: 10.3390/jtaer20020102.
C. Kuchler, A. Stoll, M. Ziegele, and T. K. Naab, “Gender-related differences in online comment sections: Findings from a large-scale content analysis of commenting behavior,” Social Science Computer Review, vol. 41, no. 3, pp. 728-747, 2023, doi: 10.1177/08944393211052042.
X. Dong, H. Liu, N. Xi, J. Liao, and Z. Yang, “Short video marketing: what, when and how short-branded videos facilitate consumer engagement,” Internet Research, vol. 34, no. 3, pp. 1104-1128, 2024, doi: 10.1108/INTR-02-2022-0121.
P. Kumar, G. Bhatt, O. Ingle, D. Goyal, and B. Raman, “Affective Feedback Synthesis Towards Multimodal Text and Image Data,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 19, no. 6, pp. 1-23, 2023, doi: 10.1145/3589186.
S. Wang, Y. Zhai, G. Xu, and N. Wang, “Multi-modal affective computing: An application in teaching evaluation based on combined processing of texts and images,” Traitement du Signal, vol. 40, no. 2, pp. 533-542, 2023, doi: 10.18280/ts.400212.
M. Firdaus, G. V. Singh, A. Ekbal, and P. Bhattacharyya, “Affect-GCN: a multimodal graph convolutional network for multiemotion with intensity recognition and sentiment analysis in dialogues,” Multimedia Tools and Applications, vol. 82, no. 28, pp. 43251-43272, 2023, doi: 10.1007/s11042-023-14885-1.
B. Schuller, A. Mallol-Ragolta, A. P. Almansa, I. Tsangko, M. M. Amin, A. Semertzidou, and S. Amiriparian, “Affective computing has changed: The foundation model disruption,” npj Artificial Intelligence, vol. 2, no. 1, pp. 16-26, 2026, doi: 10.1038/s44387-025-00061-3.
E. Zhang, H. Zong, X. Li, M. Feng, and J. Ren, “ICSF: Integrating inter-modal and cross-modal learning framework for self-supervised heterogeneous change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, no. 1, pp. 1-16, 2024, doi: 10.1109/TGRS.2024.3519195.
T. J. Wen, C. H. Chuan, G. Anghelcev, S. Sar, J. T. Yun, and Y. Xu, “Infusing affective computing models into advertising research on emotions,” Journal of advertising, vol. 53, no. 5, pp. 710-731, 2024, doi: 10.1080/00913367.2024.2409254.
Z. Li, L. Zhang, K. Zhang, Y. Zhang, and Z. Mao, “Improving image-text matching with bidirectional consistency of crossmodal alignment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 6590-6607, 2024, doi: 10.1109/TCSVT.2024.3369656.
A. Zavras, D. Michail, B. Demir, and I. Papoutsis, “Mind the modality gap: Towards a remote sensing vision-language model via cross-modal alignment,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 228, no. 1, pp. 270-287, 2025, doi: 10.1016/j.isprsjprs.2025.06.019.
M. Lan, F. Rong, H. Jiao, Z. Gao, and L. Zhang, “Language query-based transformer with multiscale cross-modal alignment for visual grounding on remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, no. 1, pp. 1-13, 2024, doi: 10.1109/TGRS.2024.3407598.
Y. Ye, J. Zhang, Z. Chen, and Y. Xia, “CADS: A self-supervised learner via cross-modal alignment and deep self-distillation for CT volume segmentation,” IEEE Transactions on Medical Imaging, vol. 44, no. 1, pp. 118-129, 2024, doi: 10.1109/TMI.2024.3431916.
L. Chen, Z. Deng, L. Liu, and S. Yin, “Multilevel semantic interaction alignment for video–text cross-modal retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 6559-6575, 2024, doi: 10.1109/TCSVT.2024.3360530.
L. Ding, L. Liu, Y. Huang, C. Li, C. Zhang, W. Wang, and L. Wang, “Text-to-image vehicle re-identification: Multi-scale multi-view cross-modal alignment network and a unified benchmark,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 7, pp. 7673-7686, 2024, doi: 10.1109/TITS.2023.3348599.
F. Bi, T. He, Y. Xie, and X. Luo, “Two-stream graph convolutional network-incorporated latent feature analysis,” IEEE Transactions on Services Computing, vol. 16, no. 4, pp. 3027-3042, 2023, doi: 10.1109/TSC.2023.3241659.
C. Cao, Y. Lu, and Y. Zhang, “Context recovery and knowledge retrieval: A novel two-stream framework for video anomaly detection,” IEEE Transactions on Image Processing, vol. 33, no. 1, pp. 1810-1825, 2024, doi: 10.1109/TIP.2024.3372466.
K. Vo, S. Truong, K. Yamazaki, B. Raj, M. T. Tran, and N. Le, “Aoe-net: Entities interactions modeling with adaptive attention mechanism for temporal action proposals generation,” International Journal of Computer Vision, vol. 131, no. 1, pp. 302-323, 2023, doi: 10.1007/s11263-022-01702-9.
S. Park, M. Mark, B. Park, and H. Hong, “Using speaker-specific emotion representations in wav2vec 2.0-based modules for speech emotion recognition,” Computers, Materials and Continua, vol. 77, no. 1, pp. 1009-1030, 2023, doi: 10.32604/cmc.2023.041332.
N. Wang and Q. Wang, “Dynamic weighted gating for enhanced cross-modal interaction in multimodal sentiment analysis,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 21, no. 1, pp. 1-19, 2024, doi: 10.1145/3702996.
H. Kumar and M. Aruldoss, “Gated Cross-Modal Fusion Mechanism for Audio-Video-based Emotion Recognition,” Engineering, Technology & Applied Science Research, vol. 15, no. 2, pp. 20835-20841, 2025, doi: 10.48084/etasr.9430.
A. G. Mengara Mengara and Y. K. Moon, “CAG-MoE: Multimodal emotion recognition with cross-attention gated mixture of experts,” Mathematics, vol. 13, no. 12, pp. 1907-1944, 2025, doi: 10.3390/math13121907.
S. Zhang, J. Liu, Y. Jiao, Y. Zhang, L. Chen, and K. Li, “A multimodal semantic fusion network with cross-modal alignment for multimodal sentiment analysis,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 21, no. 10, pp. 1-22, 2025, doi: 10.1145/3744648.
R. G. Praveen and J. Alam, “Incongruity-aware cross-modal attention for audio-visual fusion in dimensional emotion recognition,” IEEE Journal of Selected Topics in Signal Processing, vol. 18, no. 3, pp. 444-458, 2024, doi: 10.1109/JSTSP.2024.3422823.
B. Yu, C. Li, and Z. Shi, “Multi-grained feature gating fusion network for multimodal sentiment analysis: B,” Yu et al. Knowledge and Information Systems, vol. 67, no. 8, pp. 6879-6905, 2025, doi: 10.1007/s10115-025-02446-x.
R. Ji, J. Li, L. Zhang, J. Liu, and Y. Wu, “Dual transformer with multi-grained assembly for fine-grained visual classification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 9, pp. 5009-5021, 2023, doi: 10.1109/TCSVT.2023.3248791.
B. Zhong, “Fine-grained sentiment analysis using multidimensional feature fusion and GCN,” Journal of Information and Telecommunication, vol. 9, no. 1, pp. 91-112, 2025, doi: 10.1080/24751839.2024.2386785.
Y. Cui, C. Han, and D. Liu, “Collaborative multi-task learning for multi-object tracking and segmentation,” Journal on Autonomous Transportation Systems, vol. 1, no. 2, pp. 1-23, 2024, doi: 10.1145/3632181.
X. Hao, S. Yang, R. Liu, Z. Feng, T. Peng, and B. Huang, “SMTC-CL: Continuous learning via selective multi-task coordination for adaptive signal classification,” IEEE Transactions on Cognitive Communications and Networking, vol. 11, no. 3, pp. 1664-1681, 2024, doi: 10.1109/TCCN.2024.3485083. Y. Gu