Introducing the EG3D Generative Model to Build a High-Fidelity Dynamic Face Actuation System for Gamified Film and Television Content
Main Article Content
Abstract
With the increasing demand for immersion and interactivity in gamified film and television content, maintaining high-fidelity dynamic facial expressions remains challenging due to insufficient detail preservation and temporal inconsistencies across consecutive frames. To address these issues, this paper proposes EG3D-TA, a framework that integrates EG3D implicit field modeling with a Temporal Awareness (TA) actuation network. The proposed method constructs a three-dimensional generative representation based on EG3D, where Neural Radiance Fields (NeRF) implicitly model high-resolution facial geometry, while a keypoint-driven mapping network transforms two-dimensional facial expressions into continuous W+ latent encodings. An optical flow-guided latent interpolation strategy based on PWC-Net is further introduced to optimize inter-frame trajectories and improve temporal coherence. In addition, model pruning and feature caching mechanisms enable low-latency real-time inference. Experimental results demonstrate that the proposed approach achieves high visual fidelity (LPIPS: 0.12– 0.19; PSNR: 36.0–38.2 dB) and excellent motion consistency (TVD: 0.08–0.12; FVD: 78.3–88.9), while maintaining an interactive performance of 21.6 FPS at 5122 resolution with an end-to-end latency of 46.3 ms. Beyond high-fidelity digital character generation, the proposed geometry-aware and temporally consistent representation framework provides valuable methodological support for three-dimensional signal reconstruction, immersive visual communication, and intelligent sensing systems in advanced electromagnetic information processing and future integrated communication–perception applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
A. Melnik, M. Miasayedzenkau, D. Makaravets, D. Pirshtuk, E. Akbulut, D. Holzmann, et al., “Face generation and editing with stylegan: A survey,” IEEE Transactions on pattern analysis and machine intelligence, vol. 46, no. 5, pp. 3557-3576, 2024, doi: 10.1109/TPAMI.2024.3350004.
X. Tu, Y. Zou, J. Zhao, W. Ai, J. Dong, Y. Yao, et al., “Image-to-video generation via 3D facial dynamics,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 4, pp. 1805-1819, 2021, doi: 10.1109/TCSVT.2021.3083257.
F. Liu, D. Chen, F. Wang, Z. Li, and F. Xu, “Deep learning based single sample face recognition: a survey,” Artificial Intelligence Review, vol. 56, no. 3, pp. 2723-2748, 2023, doi: 10.1007/s10462-022-10240-2.
Y. Mei, W. Wang, X. Liu, W. Yong, W. Wu, Y. Zhu, et al., “Face animation based on multiple sources and perspective alignment,” Virtual Reality & Intelligent Hardware, vol. 6, no. 3, pp. 252-266, 2024, doi: 10.1016/j.vrih.2024.04.002.
M. Habermann, L. Liu, W. Xu, G. Pons-Moll, M. Zollhoefer, and C. Theobalt, “Hdhumans: A hybrid approach for high-fidelity digital humans,” Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 6, no. 3, pp. 1-23, 2023, doi: 10.1145/3606927.
K. Sun, S. Wu, N. Zhang, Z. Huang, Q. Wang, and H. Li, “Cgof++: Controllable 3d face synthesis with conditional generative occupancy fields,” IEEE transactions on pattern analysis and machine intelligence, pp. 46(2) 913-926, 2023, doi: 10.1109/TPAMI.2023.3328912.
J. Ling, X. Tan, L. Chen, R. Li, Y. Zhang, S. Zhao, et al., “StableFace: Analyzing and improving motion stability for talking face generation,” IEEE Journal of Selected Topics in Signal Processing, vol. 17, no. 6, pp. 1232-1247, 2023, doi: 10.1109/JSTSP.2023.3333552.
S. Bi, S. Lombardi, S. Saito, T. Simon, S. E. Wei, K. Mcphail, et al., “Deep relightable appearance models for animatable faces,” ACM Transactions on Graphics (ToG), vol. 40, no. 4, pp. 1-15, 2021, doi: 10.1145/3450626.3459829.
B. Chen, Z. Wang, B. Li, S. Wang, and Y. Ye, “Compact temporal trajectory representation for talking face video compression,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 11, pp. 7009-7023, 2023, doi: 10.1109/TCSVT.2023.3271130.
A. Kammoun, R. Slama, H. Tabia, T. Ouni, and M. Abid, “Generative adversarial networks for face generation: A survey,” ACM Computing Surveys, vol. 55, no. 5, pp. 1-37, 2022, doi: 10.1145/3527850.
K. Teotia, R. MB, X. Pan, H. Kim, P. Garrido, and M. Elgharib, “Hq3davatar: High-quality implicit 3d head avatar,” ACM Transactions on Graphics, vol. 43, no. 3, pp. 1-24, 2024, doi: 10.1145/3649889.
K. Singla and P. Nand, “Optimizing deep learning architectures for novel view synthesis: Investigating the impact of NeRF MLP parameters on complex scenes,” International Journal of Information Technology, vol. 16, no. 4, pp. 2295-2305, 2024, doi: 10.1007/s41870-023-01470-w.
X. Gao, C. Zhong, J. Xiang, Y. Hong, Y. Guo, and J. Zhang, “Reconstructing personalized semantic facial nerf models from monocular video,” ACM Transactions on Graphics (TOG), vol. 41, no. 6, pp. 1-12, 2022, doi: 10.1145/3550454.3555501.
K. Wang, S. Peng, X. Zhou, J. Yang, and G. Zhang, “NerfCap: Human performance capture with dynamic neural radiance fields,” IEEE Transactions on Visualization and Computer Graphics, vol. 29, no. 12, pp. 5097-5110, 2022, doi: 10.1109/TVCG.2022.3202503.
W. Y. Zhou, L. Yuan, S. Y. Chen, L. Gao, and S. M. Hu, “LC-NeRF: Local controllable face generation in neural radiance field,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 8, pp. 5437-5448, 2023, doi: 10.1109/TVCG.2023.3293653.
Y. J. Yuan, X. Han, Y. He, F. L. Zhang, and L. Gao, “Munerf: Robust makeup transfer in neural radiance fields,” IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 3, pp. 1746-1757, 2024, doi: 10.1109/TVCG.2024.3368443.
D. Han, J. Ryu, S. Kim, S. Kim, J. Park, and H. J. Yoo, “MetaVRain: A mobile neural 3-D rendering processor with bundle-frame-familiarity-based NeRF acceleration and hybrid DNN computing,” IEEE Journal of Solid-State Circuits, vol. 59, no. 1, pp. 65-78, 2023, doi: 10.1109/JSSC.2023.3291871.