Research on Animation Generation Method of Virtual Digital Human Based on Motion Capture and AI-Driven Technology
Main Article Content
Abstract
Virtual digital human animation generation faces persistent challenges in motion fidelity, facial expressiveness, and whole-body coordination. A hybrid framework integrating optical multi-camera motion capture with AI-driven generation models is proposed to address these limitations. The system employs a 16-camera optical capture array calibrated to sub-millimeter precision, combined with a Bidirectional Long Short-Term Memory (BiLSTM) network for motion sequence generation and an audio-semantic fusion module for facial animation synthesis. Experimental results demonstrate that the proposed framework achieves a mean joint position error (MPJPE) of 18.3 mm, a Fréchet Inception Distance (FID) score of 6.72 in motion generation quality, and a lip-synchronization accuracy of 94.6%. Compared with existing single-modal methods, the proposed approach reduces temporal jitter by 37.4% and improves rendering frame stability to 58.2 FPS under real-time conditions. The results confirm that the combined capture-and-generation pipeline effectively improves animation realism and computational efficiency, offering a scalable solution for virtual human deployment across entertainment, education, and interactive media domains.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
L. Chang and L. Zheng, “Human Motion Analysis and Generation Based on Optical Sensors and Deep Learning,” International Journal of Humanoid Robotics, prepublish, 2026, doi: 10.1142/S0219843626400116.
Y. Zhou, Z. Zhang, F. Wen, et al., “An All-in-One Quality Assessment Agent for 4D digital human: Bridging talking heads and animated human,” Information Processing and Management, vol. 63, no. 7PA, pp. 104817– 104817, 2026, doi: 10.1016/J.IPM.2026.104817.
Y. Wang, W. He, Q. Yao, et al., “MSadTalker: Modified Stylized Audio-Driven Single Image Talking Face Animation Based on Head Motion Generation and Visual Silence Detection,” Innovative Applications of AI, vol. 3, no. 1, pp. 30–38, 2026, doi: 10.70695/IAAI202601A8.
Z. Li, “Biomechanical analysis of cinematic motion: AI-driven generation and evaluation in film and animation,” Journal of Computational Methods in Sciences and Engineering, vol. 25, no. 6, pp. 5375–5387, 2025, doi: 10.1177/14727978251348639.
F. Christian, A. Christoph, K. Philipp, et al., “Holistic animation and functional solution for digital manikins in context of virtual assembly simulations,” Procedia CIRP, vol. 97, pp. 15–20, 2021, doi: 10.1016/J. PROCIR.2020.05.198.
M. Kim, T. Kim, and T. K. Lee, “3D Digital Human Generation from a Single Image Using Generative AI with Real-Time Motion Synchronization,” Electronics, vol. 14, no. 4, pp. 777–777, 2025, doi: 10.3390/ELEC TRONICS14040777.
Y. Huanghao, L. Jiacheng, C. Xiaohong, et al., “WeAnimate: Motion-coherent animation generation from video data,” Multimedia Tools and Applications, vol. 81, no. 15, pp. 20685–20703, 2022, doi: 10.1007/S110 42-022-12359-4.
S. Ding, “Fusion of motion smoothing algorithm and motion segmentation algorithm for human animation generation,” PloS one, vol. 20, no. 2, p. e0318979, 2025, doi: 10.1371/JOURNAL.PONE.0318979.
Y. Zhang, R. Su, J. Yu, et al., “3D facial modeling, animation, and rendering for digital humans: A survey,” Neurocomputing, vol. 598, pp. 128168–128168, 2024, doi: 10.1016/J.NEUCOM.2024.128168.
Y. Benbelkheir, A. Lerga, and O. Ardaiz, “A Virtual Reality Direct-Manipulation Tool for Posing and Animation of Digital Human Bodies: An Evaluation of Creativity Support,” Multimodal Technologies and Interaction, vol. 8, no. 7, pp. 60–60, 2024, doi: 10.3390/MTI8070060.
C. Lan, Y. Wang, C. Wang, et al., “Application of ChatGPT-Based Digital Human in Animation Creation,” Future Internet, vol. 15, no. 9, p. 300, 2023, doi: 10.3390/FI15090300.
X. Ling, Y. Zhu, W. Liu, et al., “The Generation of Articulatory Animations Based on Keypoint Detection and Motion Transfer Combined with Image Style Transfer,” Computers, vol. 12, no. 8, p. 150, 2023, doi: 10.3390/COMPUTERS12080150.
M. S. A. Abrar, S. K. Nishat, S. M. Muhammad, et al., “Deep Learning-Based Motion Style Transfer Tools, Techniques and Future Challenges,” Sensors, vol. 23, no. 5, pp. 2597–2597, 2023, doi: 10.3390/S23052597.
K. Ryotaro and K. Seiichiro, “A generative model of calligraphy based on image and human motion,” Precision Engineering, vol. 77, pp. 340–348, 2022, doi: 10.1016/J.PRECISIONENG.2022.06.006.
B. Anthony, G. RyanRhys, G. Robert, et al., “Generative model-enhanced human motion prediction,” Applied AI Letters, vol. 3, no. 2, pp. e63–e63, 2022, doi: 10.1002/AIL2.63.
B. Q. Zhao, Y. Y. Fu, Z. Su, et al., “Review on 3D digital human motion generation guided by multimodal information,” Journal of Image and Graphics, vol. 29, no. 9, pp. 2541–2565, 2024.