Research on an AI Visual Effects Generation Framework for Virtual Digital Humans in Film and Television Production

Main Article Content

P. Yan
Y. B. Ma
H. L. Li
Q. Li
J. Zhang

Abstract

The large-scale application of virtual digital humans in film and television has imposed extremely high demands on the generation precision and systematicity of visual effects. To address three major technical bottlenecks—high-fidelity 3D reconstruction of film-grade virtual digital humans, temporal driving consistency, and rendering realism in complex scenes—this paper presents a modular AI visual effects generation framework tailored to film and television production pipelines. The framework achieves collaborative geometry-texture modeling through joint optimization of NeRF and GAN, and resolves stage-wise representation inconsistency via cross-module feature alignment. Leveraging Transformer-based multi-scale temporal modeling and joint control of skeletal keypoints, high-precision motion and expression restoration is realized. A dual-path rendering pipeline combining physical rendering and diffusion models ensures cross-scene visual consistency. It also supports multi-level quality adaptation to meet the differentiated demands of preview, editing and final delivery in actual production workflows. Experiments demonstrate that the framework effectively improves core metrics of digital human production, balances computational efficiency and output quality, and demonstrates practical viability for industrial deployment in film and television.

Downloads

Download data is not yet available.

Article Details

How to Cite
Yan, P., Ma, Y. B., Li, H. L., Li, Q., & Zhang, J. (2026). Research on an AI Visual Effects Generation Framework for Virtual Digital Humans in Film and Television Production. Advanced Electromagnetics, 15(3), 10837–10845. https://doi.org/10.7716/aem.v15i3.4291
Section
Research Articles

References

B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: representing scenes as neural radiance fields for view synthesis,” Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2022, doi: 10.1145/3503250.

View Article

Z. Ye, Z. Jiang, Y. Ren, et al., “GeneFace: generalized and high-fidelity audio-driven 3D talking face synthesis,” in Proc. International Conference on Learning Representations, Kigali, Rwanda, May 1–5, 2023. Kigali, Rwanda: OpenReview, 2023.

O. Gafni, J. Thies, M. Zollhöfer, et al., “Dynamic neural radiance fields for monocular 4D facial avatar reconstruction,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, Jun. 19–25, 2021. Piscataway, NJ: IEEE, pp. 8649–8658, 2021.

S. Qian, U. Kirschstein, S. Davoli, et al., “GaussianAvatars: photorealistic head avatars with rigged 3D Gaussians,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, Jun. 17– 21, 2024. Piscataway, NJ: IEEE, pp. 4789–4799, 2024.

Z. Ye, J. He, Z. Jiang, et al., “GeneFace++: generalized and stable real-time audio-driven 3D talking face generation,” arXiv, 2023. [Online]. Available: https://arxiv.org/abs/2305.00787

View Article

Y. Zheng, Y. Wenzheng, W. Chen, et al., “IM Avatar: implicit morphable head avatars from videos,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, Jun. 18–24, 2022. Piscataway, NJ: IEEE, pp. 13545–13555, 2022.

H. Yang, H. Zhu, Y. Wang, et al., “EnNeRFACE: neural radiance fields for face animation and reenactment,” The Visual Computer, vol. 39, no. 5, pp. 1865–1878, 2023.

T. Jiang, X. Chen, J. Song, and O. Hilliges, “InstantAvatar: learning avatars from monocular video in 60 seconds,” in Proc. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, Jun. 18–22, 2023. Piscataway, NJ: IEEE, pp. 16922–16932, 2023, doi: 10.1109/CVPR52729.2023.01623.

View Article

J. Wang, J.-C. Xie, X. Li, F. Xu, C.-M. Pun, and H. Gao, “Gaussian-Head: high-fidelity head avatars with learnable Gaussian derivation,” IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 7, pp. 4141–4154, 2025, doi: 10.1109/TVCG.2025.3561794.

View Article

J. J. Park, P. R. Florence, J. Straub, R. A. Newcombe, and S. Lovegrove, “DeepSDF: learning continuous signed distance functions for shape representation,” in Proc. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, Jun. 15–20, 2019. Piscataway, NJ: IEEE, pp. 165–174, 2019.

J. Liang, Y. Liu, K. Wang, et al., “Perceptual quality assessment for novel view synthesis,” Computer Graphics Forum, vol. 43, no. 1, e15036, 2024.

W. Zhang, et al., “SadTalker: learning realistic 3D motion coefficients for stylized audio-driven single image talking face animation,” in Proc. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, Jun. 18–22, 2023. Piscataway, NJ: IEEE, pp. 8652–8661, 2023, doi: 10.1109/CVPR52729.2023.00836.

View Article

Z. Sun, T. Lv, S. Ye, M. Lin, J. Sheng, Y. Wen, M. Yu, and Y. Liu, “DiffPoseTalk: speech-driven stylistic 3D facial animation and head pose generation via diffusion models,” ACM Transactions on Graphics, vol. 43, pp. 1–9, 2023.

S. Saito, T. Simon, J. Saragih, et al., “Relightable Gaussian codec avatars,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, Jun. 17–21, 2024. Piscataway, NJ: IEEE, pp. 130–139, 2024.

J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. 34th International Conference on Neural Information Processing Systems (NIPS ’20), Virtual, Dec. 6–12, 2020. Red Hook, NY, USA: Curran Associates, pp. 6840–6851, 2020.

Y. X. Pan, “Research on Key Technologies of Metaverse Virtual Space Construction Based on Blockchain and Deep Learning,” Ph.D. dissertation, Beijing University of Posts and Telecommunications, Beijing, 2025, doi: 10.26969/d.cnki.gbydu.2025.000150.

View Article

Most read articles by the same author(s)

1 2 > >>