XR Immersive Film and Television Entertainment Interactive Narrative System Based on Multimodal Perception
Main Article Content
Abstract
Extended Reality (XR) immersive entertainment systems increasingly rely on real-time multimodal sensing and interactive information transmission to achieve adaptive user experiences. However, existing XR narrative frameworks are primarily constrained by single-modal interaction and limited contextual awareness, resulting in insufficient personalization and delayed narrative adaptation. This study proposes a multimodal perception-driven interactive narrative system that integrates visual streams, speech signals, and physiological sensing through a unified acquisition and temporal synchronization framework. A Transformer-based cross-modal fusion network is employed to jointly encode heterogeneous sensory information and infer user emotions and behavioral intentions, while a reinforcement learning policy dynamically optimizes narrative evolution and cooperates with a large language model for context-aware content generation. An immersive XR feedback loop further coordinates visual rendering, spatial audio propagation, and haptic interaction to establish continuous perception–decision–feedback cycles. Experimental results demonstrate that multimodal emotion recognition maintains an accuracy above 0.80, intention recognition reaches 0.91 during later interaction stages, and users’ subjective immersion scores increase from 3.5 to 6.2 with stable low-latency performance. By coupling multimodal sensing, spatial signal propagation, and adaptive interaction within an integrated XR architecture, the proposed framework provides an engineering-oriented paradigm for electromagnetic sensing-enabled immersive environments, intelligent wireless perception, and next-generation human–machine communication systems.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
H. Wang, Z. Feng, X. Yang, et al., “MRLab: Virtual-reality fusion smart laboratory based on multimodal fusion,” International Journal of Human-Computer Interaction, vol. 40, no. 8, pp. 1975-1988, 2024, doi: 10.1080/10447318.2023.2227823.
C. Chen, K.ZK. Zhang, Z. Chu, et al., “Augmented reality in the metaverse market: the role of multimodal sensory interaction,” Internet Research, vol. 34, no. 1, pp. 9-38, 2024, doi: 10.1108/INTR-08-2022-0670.
B. Hussain, J. Guo, F. Sidra, et al., “Enhancing spatial awareness via multimodal fusion of cnn-based visual and depth features,” International Journal of Ethical AI Application, vol. 1, no. 3, pp. 13-27, 2025, doi: 10.64229/gdz8tc37.
L. Cao, H. Zhang, C. Peng, et al., “Real-time multimodal interaction in virtual reality-a case study with a large virtu-al interface,” Multimedia Tools and Applications, vol. 82, no. 16, pp. 25427-25448, 2023, doi: 10.1007/s11042-023-14381-6.
J.A. Duncan, F. Alambeigi, and W. Pryor M, “A survey of multimodal perception methods for human-robot interaction in social environments,” ACM Transactions on Human-Robot Interaction, vol. 13, no. 4, pp. 1-50, 2024, doi: 10.1145/3657030.
J. Zhang, S. Wang, W. He, et al., “Perception and Decision-Making for Multi Modal Interaction Based on Fuzzy Theory in the Dynamic Environment,” International Journal of Human-Computer Interaction, vol. 40, no. 24, pp. 8794-8808, 2024, doi: 10.1080/10447318.2023.2291607.
A. Nurjamin, L. R. Nurjamin, Y.N. Fajriah, et al., “Developing and evaluating an augmented reality (AR) digital storytelling video to foster multimodal literacy and narrative comprehension,” Journal of Engineering Science and Technology, vol. 20, no. 4, pp. 919-956, 2025.
J. Shao and D. Wu, “Evaluation on algorithms and models for multimodal information fusion and evaluation in new media art and film and television cultural creation,” Journal of Computational Methods in Science and Engineering, vol. 24, no. 4-5, pp. 3173-3189, 2024, doi: 10.3233/JCM-247565.
Y. Ji, “Artistic alchemy: Exploring the fusion of art theory and film aesthetics in visual storytelling,” Herança, vol. 7, no. 2, pp. 51-68, 2024, doi: 10.52152/heranca.v7i2.785.
S. M. Jamasbi and A. Ghazvineh, “A Multimodal Analysis of Films: Toward a Peircean Framework,” Quarterly Review of Film and Video, vol. 42, no. 6, pp. 1543-1565, 2025, doi: 10.1080/10509208.2023.2265783.
Z. Cai and K. Liu, “Construction of interactive narrative in children’s drama driven by generative adversarial networks,” International Journal of Information and Communication Technology, vol. 27, no. 30, pp. 1-23, 2026, doi: 10.1504/IJICT.2026.152655.
A. Raheel, D. Khalid, and S. S. Ahmed, “Classifying Emotions in 3-D: Physiological Insights Into Tactile-Audio-Visual Immersion,” IEEE Sensors Journal, vol. 26, no. 2, pp. 2535-2542, 2025, doi: 10.1109/JSEN.2025.3637217.
P. Qi, “Movie visual and speech analysis through multimodal llm for recommendation systems,” IEEE Access, vol. 12, no. 1, pp. 145686-145702, 2024, doi: 10.1109/ACCESS.2024.3471568.
Y. Liu and J. Li, “" Expansion" and" Obstacles": The New Wave of Intelligent Media Empowering the Development of Film and Television Arts with Digital Technology,” Journal of Social Science Humanities and Literature, vol. 7, no. 6, pp. 57-66, 2024, doi: 10.53469/jsshl.2024.07(06).11.
Y. Zou, “The Value Logic, Challenges, and Development Strategies of Generative AI Empowering Film and Television Production,” Advances in Education, Humanities and Social Science Research, vol. 15, no. 1, pp. 662-662, 2025, doi: 10.56028/aehssr.15.1.662.2025.
X. Zhu, C. Guo, H. Feng, et al., “A review of key technologies for emotion analysis using multimodal information,” Cognitive Computation, vol. 16, no. 4, pp. 1504-1530, 2024, doi: 10.1007/s12559-024-10287-z.
C. Yang, “Application Scenarios and Creation Paradigms of Artificial Intelligence in Digital Media Design,” Computer Life, vol. 14, no. 1, pp. 50-53, 2026, doi: 10.54097/qsz2df37.
C. Zhang and Q. Lei, “Intersemiotic translation of The Song of Everlasting Sorrow from narrative poetry to dance drama: a multimodal stylistics perspective,” Semiotica, vol. 2026, no. 268, pp. 95-128, 2026, doi: 10.1515/sem-2024-0213.
L. Zhang and J. Wang, “Hot Topics and Frontier Evolution in Music Research within the Multimodal Art Context: A Bibliometric Analysis Based on CiteSpace,” Korean Science and Art Forum, vol. 44, no. 1, pp. 619-639, 2026, doi: 10.17548/ksaf.2026.01.30.619.
J. Yi, Y. Tian, and Y. Zhao, “Design of red culture retrieval system based on multimodal data fusion and innovation of communication strategy path,” IEEE Access, vol. 11, no. 1, pp. 134118-134125, 2023, doi: 10.1109/ACCESS.2023.3336419.
Z. Su, S. Lin, L. Zhang, et al., “Multitask Learning-Based Affective Prediction for Videos of Films and TV Scenes,” Applied Sciences, vol. 14, no. 11, p. 4392, 2024, doi: 10.3390/app14114391.
J. Chen, K.P. Seng, J. Smith, et al., “Situation awareness in ai-based technologies and multimodal systems: Architectures, challenges and applications,” IEEE Access, vol. 12, no. 1, pp. 88779-88818, 2024, doi: 10.1109/ACCESS.2024.3416370.
M. Dokoupil, “Experiencing Landscape as a Visual Event: Dynamic Hyperstereoscopy and the Emergence of Perceptual Space,” International Journal on Stereo & Immersive Media, vol. 9, no. 2, pp. 120-139, 2025, doi: 10.24140/ijsim.v9i2.9445.
W. Yuan, J. Chen, S. Chen, et al., “Transformer in reinforcement learning for decision-making: a survey,” Frontiers of Information Technology & Electronic Engineering, vol. 25, no. 6, pp. 763-790, 2024, doi: 10.1631/FITEE.2300548.
S. Hu, L. Shen, Y. Zhang, et al., “On transforming reinforcement learning with transformers: The development trajectory,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 8580-8599, 2024, doi: 10.1109/TPAMI.2024.3408271.
T. Chen and L. Mo, “Swin-fusion: swin-transformer with feature fusion for human action recognition,” Neural Processing Letters, vol. 55, no. 8, pp. 11109-11130, 2023, doi: 10.1007/s11063-023-11367-1.
L. Chen, H. Zhao, C. Shi, et al., “Enhancing multimodal perception and interaction: An augmented reality visualization system for complex decision making,” Systems, vol. 12, no. 1, pp. 7-9, 2023, doi: 10.3390/systems12010007.
Y. Wang, M. Guizani, and S. Hossain M, “Artificial Intelligence for Virtual Reality: State of the Art, Challenges, and Future Perspectives,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 21, no. 11, pp. 1-29, 2025, doi: 10.1145/3769090.
R. Abdulrahman, A. Jamil, A. Amjad, et al., “Automated Deep Learning Approaches for Multimodal Emotion Recognition: A Review of Fusion Strategies, Modalities and Architectures,” Machines and Algorithms, vol. 4, no. 3, pp. 198-214, 2025, doi: 10.66108/mna.v4i3.103.
B. Liang, N. Li, G. Li, et al., “Sensing Technologies for Hand Gesture Recognition in Human-Robot Interaction: A Review,” IEEE Sensors Journal, vol. 26, no. 2, pp. 1501-1519, 2025, doi: 10.1109/JSEN.2025.3635622.
C. Mao, “An Intelligent Decision-Support Framework Using Mobile and Virtual Reality Technologies for Optimising Intangible Cultural Heritage Management,” Decision Making: Applications in Management and Engineering, vol. 8, no. 2, pp. 881-897, 2025, doi: 10.31181/dmame8220251631.
D. Wu, Z. Yang, P. Zhang, et al., “Virtual-reality interpromotion technology for metaverse: A survey,” IEEE Internet of Things Journal, vol. 10, no. 18, pp. 15788-15809, 2023, doi: 10.1109/JIOT.2023.3265848.
J. Ding, Y. Zhang, Y. Shang, et al., “Understanding world or predicting future? a comprehensive survey of world models,” ACM Computing Surveys, vol. 58, no. 3, pp. 1-38, 2025, doi: 10.1145/3746449. J. Shang