Optimizing Digital Art Style Transfer Using a Diffusion Model
Main Article Content
Abstract
Current generative models are prone to content distortion when guided by strong digital art styles, limiting the preservation of structural and semantic information during style transfer. Such controllable visual representation and semantic consistency are also important for intelligent imaging, computational perception, and electromagnetic information processing systems that require reliable feature preservation across heterogeneous visual domains. To address this issue, this study proposes a latent diffusion model (LDM)-based framework for optimized digital art style transfer. Strong style is quantified by the normalized norm of the CLIP-encoded style embedding in the latent space, and its distribution is employed as a continuous indicator to regulate style injection during diffusion. A Cross-Attention mechanism is introduced into the inverse denoising process to dynamically fuse content and style features while preventing structural distortion caused by excessive style guidance. In addition, a linear strength factor enables continuous adjustment of style intensity, and a pre-trained ControlNet generates regional guidance masks to spatially constrain style propagation in the latent space. Guided by style encoding, the predicted noise is iteratively refined under a fixed diffusion schedule to recover high-quality latent representations while preserving semantic consistency. Experiments on the WikiArt and COCO datasets demonstrate that the proposed method effectively alleviates content distortion and achieves a PSNR of 27.6 ± 1.2 dB and an SSIM of 0.85 ± 0.02. The proposed framework provides an effective solution for controllable image generation and offers potential value for computational imaging, intelligent visual sensing, and electromagnetic information processing applications requiring high-fidelity semantic representation.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
L. Yu and Q. Zheng, “AI-Enhanced Digital Creativity Design: Content-Style Alignment for Image Stylization,” IEEE Access, vol. 11, pp. 143964-143979, 2023, doi: 10.1109/ACCESS.2023.3342444.
X. Han, Y. Wu, and R. Wan, “A method for style transfer from artistic images based on depth extraction generative adversarial network,” Applied Sciences, vol. 13, no. 2, pp. 867-881, 2023, doi: 10.3390/app13020867.
H. Li, L. Wang, and J. Liu, “A review of deep learning-based image style transfer research,” The Imaging Science Journal, vol. 73, no. 4, pp. 504-526, 2025, doi: 10.1080/13682199.2024.2418216.
Y. Xu, M. Xia, K. Hu, et al., “Style Transfer Review: Traditional Machine Learning to Deep Learning,” Information, vol. 16, no. 2, pp. 157-220, 2025, doi: 10.3390/info16020157.
S. Gayatri, G. Reddy S, T. Sirisha, et al., “SC-GAN: A Style-Conditioned Generative Adversarial Network for High-Quality Artistic Image Generation,” Macaw International Journal of Advanced Research in Computer Science and Engineering, vol. 11, no. 1, pp. 1-10, 2025.
Y. Zhang, B. Hu, Y. Huang, et al., “Adaptive style modulation for artistic style transfer,” Neural Processing Letters, vol. 55, no. 5, pp. 6213-6230, 2023, doi: 10.1007/s11063-022-11135-7.
M. Xu, S. Yoon, A. Fuentes, et al., “Style-consistent image translation: A novel data augmentation paradigm to improve plant disease recognition,” Frontiers in plant science, vol. 12, pp. 773142-773158, 2022, doi: 10.3389/fpls.2021.773142.
G. Song, L. Luo, J. Liu, et al., “Agilegan: stylizing portraits by inversion-consistent transfer learning,” ACM Transactions on Graphics (TOG), vol. 40, no. 4, pp. 1-13, 2021, doi: 10.1145/3450626.3459771.
Y. Zhou, K. Yu, M. Wang, et al., “Speckle noise reduction for OCT images based on image style transfer and conditional GAN,” IEEE Journal of Biomedical and Health Informatics, vol. 26, no. 1, pp. 139-150, 2021, doi: 10.1109/JBHI.2021.3074852.
R. Fernandez-Fernandez, G. Victores J, J. Gago J, et al., “Neural policy style transfer,” Cognitive Systems Research, vol. 72, pp. 23-32, 2022, doi: 10.1016/j.cogsys.2021.11.003.
H. He, X. Chen, C. Wang, et al., “Diff-font: Diffusion model for robust one-shot font generation,” International Journal of Computer Vision, vol. 132, no. 11, pp. 5372-5386, 2024, doi: 10.1007/s11263-024-02137-0.
L. Yang, J. Liu, S. Hong, et al., “Improving diffusion-based image synthesis with context prediction,” Advances in Neural Information Processing Systems, vol. 36, pp. 37636-37656, 2023, doi: 10.52202/075280-1636.
S. Zhao, D. Chen, C. Chen Y, et al., “Uni-controlnet: All-in-one control to text-to-image diffusion models,” Advances in Neural Information Processing Systems, vol. 36, pp. 11127-11150, 2023, doi: 10.52202/075280-0491.
D. Mukherkjee, P. Saha, D. Kaplun, et al., “Brain tumor image generation using an aggregation of GAN models with style transfer,” Scientific reports, vol. 12, no. 1, pp. 9141-9157, 2022, doi: 10.1038/s41598-022-12646-y.
W. Hu and Y. Zhang, “Research on Artistic Pattern Generation for Clothing Design Based on Style Transfer,” J. COMBIN. MATH. COMBIN. COMPUT, vol. 127, pp. 4539-4550, 2025, doi: 10.61091/jcmcc127b-248.
G. Castellano and G. Vessio, “Deep learning approaches to pattern extraction and recognition in paintings and drawings: An overview,” Neural Computing and Applications, vol. 33, no. 19, pp. 12263-12282, 2021, doi: 10.1007/s00521-021-05893-z.