Recognition and Modeling of User Interaction Behaviors in Virtual Reality Based on 3D Convolutional Neural Networks
Main Article Content
Abstract
Accurate recognition of continuous user interactions in virtual reality (VR) is essential for intelligent perception systems and provides methodological support for multimodal information processing in electromagnetic sensing and nextgeneration human–machine interaction. However, fragmented action sequences, weak long-range dependencies, and semantic discontinuities remain major challenges in complex immersive environments. This study proposes a crosssegment spatiotemporal attention and behavior pattern modeling framework based on a three-dimensional convolutional neural network (3D CNN). The framework integrates multi-layer 3D convolution with residual learning for robust feature extraction, employs cross-segment spatiotemporal attention and Bidirectional Long Short-Term Memory (BiLSTM) networks to capture long-range temporal dependencies, and introduces graph neural networks with boundary smoothing and graph consistency constraints to preserve behavioral structure and action continuity. Experimental results demonstrate an average recognition accuracy of 87.2% for complex actions, a behavior pattern consistency of 95.2%, and real-time inference performance of 39.2 FPS while maintaining superior robustness under occlusion and motion blur. Beyond VR interaction analysis, the proposed hierarchical spatiotemporal representation and structured modeling strategy provides useful insights for adaptive signal interpretation, multimodal perception, and intelligent decisionmaking in electromagnetic sensing and wireless interactive systems.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
C. Cheng and H. Xu, “A 3D motion image recognition model based on 3D CNN-GRU model and attention mechanism,” Image and Vision Computing, vol. 146, no. 1, pp. 104991-105003, 2024, doi: 10.1016/j.imavis.2024.104991.
M. Liu, W. Li, B. He, C. Wang, and L. Qu, “Human Action Recognition Based on 3D Convolution and Multi-Attention Transformer,” Applied Sciences, vol. 15, no. 5, pp. 2695, 2025, doi: 10.3390/app15052695.
M. Li, F. Pan, Y. Zhang, T. Chen, and H. Du, “Integrating deep learning model and virtual reality technology for motion prediction in emergencies,” Safety Science, vol. 183, no. 1, pp. 106721-106733, 2025, doi: 10.1016/j.ssci.2024.106721.
A. Kamali Mohammadzadeh, E. Alinezhad, and S. Masoud, “Neural-Network-Driven Intention Recognition for Enhanced Human–Robot Interaction: A Virtual-Reality-Driven Approach,” Machines, vol. 13, no. 5, pp. 414-434, 2025, doi: 10.3390/machines13050414.
H. Dawood, M. Nawaz, T. Nazir, A. Javed, A. K. J. Saudagar, and H. S. AlSagri, “ARNet: Integrating Spatial and Temporal Deep Learning for Robust Action Recognition in Videos,” Computer Modeling in Engineering & Sciences (CMES), vol. 144, no. 1, pp. 1-14, 2025, doi: 10.32604/cmes.2025.066415.
H. Ullah and A. Munir, “Human action representation learning using an attention-driven residual 3dcnn network,” Algorithms, vol. 16, no. 8, pp. 369-391, 2023, doi: 10.3390/a16080369.
X. Huang and Z. Cai, “A review of video action recognition based on 3D convolution,” Computers and Electrical Engineering, vol. 108, no. 1, pp. 108713-108725, 2023, doi: 10.1016/j.compeleceng.2023.108713.
P. Raimbaud, R. Lou, F. Danglade, P. Figueroa, J. T. Hernandez, and F. Merienne, “A task-centred methodology to evaluate the design of virtual reality user interactions: A case study on hazard identification,” Buildings, vol. 11, no. 7, pp. 277-299, 2021, doi: 10.3390/buildings11070277.
B. Gan, C. Zhang, Y. Chen, and Y. C. Chen, “Research on role modeling and behavior control of virtual reality animation interactive system in Internet of Things,” Journal of Real-Time Image Processing, vol. 18, no. 4, pp. 1069-1083, 2021, doi: 10.1007/s11554-020-01046-y.
M. Zong, R. Wang, Z. Chen, M. Wang, X. Wang, and J. Potgieter, “Multi-cue based 3D residual network for action recognition,” Neural Computing and Applications, vol. 33, no. 10, pp. 5167-5181, 2021, doi: 10.1007/s00521-020-05313-8.
H. Zhang, Z. Hu, D. Yu, L. Guan, X. Liu, and C. Ma, “Multipath attention and adaptive gating network for video action recognition,” Neural Processing Letters, vol. 56, no. 2, pp. 124-144, 2024, doi: 10.1007/s11063-024-11591-3.
A. Sánchez-Caballero, D. Fuentes-Jiménez, and C. Losada-Gutiérrez, “Real-time human action recognition using raw depth video-based recurrent neural networks,” Multimedia Tools and Applications, vol. 82, no. 11, pp. 16213-16235, 2023, doi: 10.1007/s11042-022-14075-5.
H. Li, X. Li, L. Su, D. Jin, J. Huang, and D. Huang, “Deep spatio-temporal adaptive 3d convolutional neural networks for traffic flow prediction,” ACM Transactions on Intelligent Systems and Technology (TIST), vol. 13, no. 2, pp. 1-21, 2022, doi: 10.1145/3510829.
G. L. Zhang, “Design of virtual reality augmented reality mobile platform and game user behavior monitoring using deep learning,” International Journal of Electrical Engineering & Education, vol. 60, no. 2_suppl, pp. 205-221, 2023, doi: 10.1177/0020720920931079.
C. Linse, H. Alshazly, and T. Martinetz, “A walk in the black-box: 3D visualization of large neural networks in virtual reality,” Neural Computing and Applications, vol. 34, no. 23, pp. 21237-21252, 2022, doi: 10.1007/s00521-022-07608-4.
H. Li, “Convolutional Neural Network-Based Virtual Reality Real-Time Interactive System Design for Unity3D,” Computational Intelligence and Neuroscience, vol. 2022, no. 1, pp. 2530836-2530848, 2022, doi: 10.1155/2022/2530836.
N. Tasnim and J. H. Baek, “Deep learning-based human action recognition with key-frames sampling using ranking methods,” Applied Sciences, vol. 12, no. 9, pp. 4165-4183, 2022, doi: 10.3390/app12094165.
S. Guan, H. Lu, L. Zhu, and G. Fang, “AFE-CNN: 3D skeleton-based action recognition with action feature enhancement,” Neurocomputing, vol. 514, no. 1, pp. 256-267, 2022, doi: 10.1016/j.neucom.2022.10.016.
Y. Ming, F. Feng, C. Li, and J. H. Xue, “3D-TDC: A 3D temporal dilation convolution framework for video action recognition,” Neurocomputing, vol. 450, no. 1, pp. 362-371, 2021, doi: 10.1016/j.neucom.2021.03.120.
E. M. Saoudi, J. Jaafari, and S. J. Andaloussi, “Advancing human action recognition: A hybrid approach using attention-based LSTM and 3D CNN,” Scientific African, vol. 21, no. 1, pp. 01796-01816, 2023, doi: 10.1016/j.sciaf.2023.e01796.
B. Chen, H. Tang, Z. Zhang, G. Tong, and B. Li, “Video-based action recognition using spurious-3D residual attention networks,” IET Image Processing, vol. 16, no. 11, pp. 3097-3111, 2022, doi: 10.1049/ipr2.12541.
M. Toshpulatov, W. Lee, S. Lee, H. Yoon, and U. Kang, “DDC3N: Doppler-driven convolutional 3D network for human action recognition,” IEEE Access, vol. 12, no. 1, pp. 93546-93567, 2024, doi: 10.1109/ACCESS.2024.3422428.
S. Pehlivan and J. Laaksonen, “Temporal teacher with masked transformers for semi-supervised action proposal generation,” Machine Vision and Applications, vol. 35, no. 3, pp. 36-51, 2024, doi: 10.1007/s00138-024-01521-7.
Y. You, L. Zhang, P. Tao, S. Liu, and L. Chen, “Spatiotemporal transformer neural network for time-series forecasting,” Entropy, vol. 24, no. 11, pp. 1651-1669, 2022, doi: 10.3390/e24111651.
N. Aziere and S. Todorovic, “Multistage temporal convolution transformer for action segmentation,” Image and Vision Computing, vol. 128, pp. 104567-104577, 2022, doi: 10.1016/j.imavis.2022.104567.
M. Li, X. Wang, H. Wang, and M. Yang, “LTGS-Net: Local Temporal and Global Spatial Network for Weakly Supervised Video Anomaly Detection,” Sensors, vol. 25, no. 16, pp. 4884-4901, 2025, doi: 10.3390/s25164884.