Research on an AI Visual Effects Generation Framework for Virtual Digital Humans in Film and Television Production
Main Article Content
Abstract
The large-scale application of virtual digital humans in film and television has imposed extremely high demands on the generation precision and systematicity of visual effects. To address three major technical bottlenecks—high-fidelity 3D reconstruction of film-grade virtual digital humans, temporal driving consistency, and rendering realism in complex scenes—this paper presents a modular AI visual effects generation framework tailored to film and television production pipelines. The framework achieves collaborative geometry-texture modeling through joint optimization of NeRF and GAN, and resolves stage-wise representation inconsistency via cross-module feature alignment. Leveraging Transformer-based multi-scale temporal modeling and joint control of skeletal keypoints, high-precision motion and expression restoration is realized. A dual-path rendering pipeline combining physical rendering and diffusion models ensures cross-scene visual consistency. It also supports multi-level quality adaptation to meet the differentiated demands of preview, editing and final delivery in actual production workflows. Experiments demonstrate that the framework effectively improves core metrics of digital human production, balances computational efficiency and output quality, and demonstrates practical viability for industrial deployment in film and television.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: representing scenes as neural radiance fields for view synthesis,” Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2022, doi: 10.1145/3503250.
Z. Ye, Z. Jiang, Y. Ren, et al., “GeneFace: generalized and high-fidelity audio-driven 3D talking face synthesis,” in Proc. International Conference on Learning Representations, Kigali, Rwanda, May 1–5, 2023. Kigali, Rwanda: OpenReview, 2023.
O. Gafni, J. Thies, M. Zollhöfer, et al., “Dynamic neural radiance fields for monocular 4D facial avatar reconstruction,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, Jun. 19–25, 2021. Piscataway, NJ: IEEE, pp. 8649–8658, 2021.
S. Qian, U. Kirschstein, S. Davoli, et al., “GaussianAvatars: photorealistic head avatars with rigged 3D Gaussians,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, Jun. 17– 21, 2024. Piscataway, NJ: IEEE, pp. 4789–4799, 2024.
Z. Ye, J. He, Z. Jiang, et al., “GeneFace++: generalized and stable real-time audio-driven 3D talking face generation,” arXiv, 2023. [Online]. Available: https://arxiv.org/abs/2305.00787
Y. Zheng, Y. Wenzheng, W. Chen, et al., “IM Avatar: implicit morphable head avatars from videos,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, Jun. 18–24, 2022. Piscataway, NJ: IEEE, pp. 13545–13555, 2022.
H. Yang, H. Zhu, Y. Wang, et al., “EnNeRFACE: neural radiance fields for face animation and reenactment,” The Visual Computer, vol. 39, no. 5, pp. 1865–1878, 2023.
T. Jiang, X. Chen, J. Song, and O. Hilliges, “InstantAvatar: learning avatars from monocular video in 60 seconds,” in Proc. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, Jun. 18–22, 2023. Piscataway, NJ: IEEE, pp. 16922–16932, 2023, doi: 10.1109/CVPR52729.2023.01623.
J. Wang, J.-C. Xie, X. Li, F. Xu, C.-M. Pun, and H. Gao, “Gaussian-Head: high-fidelity head avatars with learnable Gaussian derivation,” IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 7, pp. 4141–4154, 2025, doi: 10.1109/TVCG.2025.3561794.
J. J. Park, P. R. Florence, J. Straub, R. A. Newcombe, and S. Lovegrove, “DeepSDF: learning continuous signed distance functions for shape representation,” in Proc. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, Jun. 15–20, 2019. Piscataway, NJ: IEEE, pp. 165–174, 2019.
J. Liang, Y. Liu, K. Wang, et al., “Perceptual quality assessment for novel view synthesis,” Computer Graphics Forum, vol. 43, no. 1, e15036, 2024.
W. Zhang, et al., “SadTalker: learning realistic 3D motion coefficients for stylized audio-driven single image talking face animation,” in Proc. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, Jun. 18–22, 2023. Piscataway, NJ: IEEE, pp. 8652–8661, 2023, doi: 10.1109/CVPR52729.2023.00836.
Z. Sun, T. Lv, S. Ye, M. Lin, J. Sheng, Y. Wen, M. Yu, and Y. Liu, “DiffPoseTalk: speech-driven stylistic 3D facial animation and head pose generation via diffusion models,” ACM Transactions on Graphics, vol. 43, pp. 1–9, 2023.
S. Saito, T. Simon, J. Saragih, et al., “Relightable Gaussian codec avatars,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, Jun. 17–21, 2024. Piscataway, NJ: IEEE, pp. 130–139, 2024.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. 34th International Conference on Neural Information Processing Systems (NIPS ’20), Virtual, Dec. 6–12, 2020. Red Hook, NY, USA: Curran Associates, pp. 6840–6851, 2020.
Y. X. Pan, “Research on Key Technologies of Metaverse Virtual Space Construction Based on Blockchain and Deep Learning,” Ph.D. dissertation, Beijing University of Posts and Telecommunications, Beijing, 2025, doi: 10.26969/d.cnki.gbydu.2025.000150.