Animation of Historical Images: A Video Diffusion Generation Method Guided by Depth Estimation
Main Article Content
Abstract
The use of artificial intelligence and smart algorithms for the restoration, enhancement and creation of videos from historical images continues to grow. To tackle the problems of spatial structure drift, inter-frame flickering, and motion discontinuity in the animation of historical images, we propose a video diffusion generation model based on depth estimation. Based on the video diffusion model, the model includes the historical image preprocessing module, the single-frame depth estimation module, the spatial structure encoding module, and the temporal motion compensation module, using the depth feature as a conditional constraint in the reverse denoising generation of the model. Experimental results show that the proposed method achieves a PSNR of 28.76, an improvement of 1.28 over conventional video diffusion models; SSIM increases to 0.883, LPIPS decreases to 0.108, and FVD decreases to 219.54, with both generation quality and temporal stability outperforming the comparison methods.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
W. Ma, X. Yang, L. Jiao, et al., “Video diffusion generation: comprehensive review and open problems,” Artificial Intelligence Review, vol. 58, no. 11, p. 338, 2025.
S. Cai, D. Ceylan, M. Gadelha, et al., “Generative rendering: Controllable 4d-guided video generation with 2d diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 7611–7620.
G. Rayo and R. Tous, “Temporally Coherent Video Cartoonization for Animation Scenery Generation,” Electronics, vol. 13, no. 17, p. 3462, 2024.
Z. Xu, J. Zhang, J. H. Liew, et al., “Magicanimate: Temporally consistent human image animation using diffusion model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 1481–1490.
M. Niu, X. Cun, X. Wang, et al., “Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model,” in European conference on computer vision. Cham: Springer Nature Switzerland, 2024, pp. 111–128.
A. Gupta, L. Yu, K. Sohn, et al., “Photorealistic video generation with diffusion models,” in European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024, pp. 393–411.
Z. Gu, R. Yan, J. Lu, et al., “Diffusion as shader: 3d-aware video diffusion for versatile video generation control,” in Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, 2025, pp. 1–12.
G. Lei, C. Wang, R. Zhang, et al., “Animateanything: Consistent and controllable animation for video generation,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 27946–27956.
Z. Kuang, S. Cai, H. He, et al., “Collaborative video diffusion: Consistent multi-video generation with camera control,” Advances in Neural Information Processing Systems, vol. 37, pp. 16240–16271, 2024.
S. Chen, M. Xu, J. Ren, et al., “Gentron: Diffusion transformers for image and video generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6441–6451.
J. Xing, M. Xia, Y. Zhang, et al., “Dynamicrafter: Animating open-domain images with video diffusion priors,” in European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024, pp. 399–417.
R. Po, W. Yifan, V. Golyanik, et al., “State of the art on diffusion models for visual computing,” in Computer graphics forum, vol. 43, no. 2, 2024, Art. no. e15063.
X. Chen, Z. Liu, M. Chen, et al., “Livephoto: Real image animation with text-guided motion control,” in European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024, pp. 475–491.
S. Shen, W. Zhao, Z. Meng, et al., “Difftalk: Crafting diffusion models for generalized audio-driven portraits animation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 1982–1991.
Y. Wang, X. Chen, X. Ma, et al., “Lavie: High-quality video generation with cascaded latent diffusion models,” International Journal of Computer Vision, vol. 133, no. 5, pp. 3059–3078, 2025.
Y. Zhou, D. Zhou, M. M. Cheng, et al., “Storydiffusion: Consistent self-attention for long-range image and video generation,” Advances in Neural Information Processing Systems, vol. 37, pp. 110315–110340, 2024.
W. Xie, A. Hu, Q. Xie, et al., “Bibliometric analysis and review of AI-based video generation: research dynamics and application trends (2020– 2025),” Discover Computing, vol. 28, no. 1, p. 130, 2025.
Y. Wu, X. Deng, H. Shao, et al., “Anime Generation through Diffusion and Language Models: A Comprehensive Survey of Techniques and Trends,” Computer Modeling in Engineering & Sciences, vol. 144, no. 3, p. 2709, 2025.
Y. Jiang, C. Yu, C. Cao, et al., “Animate3d: Animating any 3d model with multi-view video diffusion,” Advances in Neural Information Processing Systems, vol. 37, pp. 125879–125906, 2024.
B. Li, C. Zheng, W. Zhu, et al., “Vivid-zoo: Multi-view video generation with diffusion model,” Advances in Neural Information Processing Systems, vol. 37, pp. 62189–62222, 2024.
W. Song, X. Wang, Y. Jiang, et al., “Expressive 3d facial animation generation based on local-to-global latent diffusion,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 11, pp. 7397–7407, 2024.
C. Ding and R. Bhowmik, “Artificial intelligence in multimedia content generation: a review of audio and video synthesis techniques,” Journal of the society for information display, vol. 34, no. 2, pp. 49–67, 2026.
H. Kandala, J. Gao, and J. Yang, “Pix2gif: Motion-guided diffusion for gif generation,” in European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024, pp. 35–51.
M. Elmoghany, R. Rossi, S. Yoon, et al., “A survey on long-video storytelling generation: architectures, consistency, and cinematic quality,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 7023–7035.
Z. Fang, K. Zhu, Z. Liu, et al., “Panoramic video generation with pre-trained diffusion models,” Advances in Neural Information Processing Systems, vol. 38, pp. 12486–12508, 2026.