Research on AI-Driven Generation Mechanisms and Design Methodologies for Digital Media Art
Main Article Content
Abstract
To address the problems of insufficient semantic control precision and lack of structured support for human-AI collaboration in digital media art generation, this study proposes a text–image cross-modal collaborative generation framework for artistic creation. Taking diffusion models and generative adversarial networks as generation kernels, the framework achieves high-precision mapping from text intent to visual features through cross-modal semantic alignment mechanisms, and realizes fine-grained control of the generation process by separating style, content, and structure variables via a conditional disentanglement module. A closed-loop human-AI collaboration process integrating intent modeling, constraint injection, and aesthetic evaluation is constructed, forming a reusable creative methodology system. Through multiple comparative experiments, the framework is quantitatively validated from generation quality, style diversity, and human-AI collaboration efficiency, demonstrating superior generation controllability and creative efficiency over single-model baseline methods. This study reveals the intrinsic mechanisms of AI involvement in digital media art creation, providing scientific methodological references for intelligent system design in this domain.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
T. Karras, M. Aittala, J. Hellsten, S. Laine, J. Lehtinen, and T. Aila, “Training Generative Adversarial Networks with Limited Data,” arXiv, abs/2006.06676, 2020.
J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” arXiv, abs/2006.11239, 2020.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” 2021, doi: 10.48550/arXiv.2112.10752.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, et al., “Learning transferable visual models from natural language supervision,” 2021, doi: 10.48550/arXiv.2103.00020.
L. Zhang, A. Rao, and M. Agrawala, “Adding Conditional Control to Text-to-Image Diffusion Models,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3813–3824, 2023.
E. J. Hu, Y. Shen, P. Wallis, A. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. International Conference on Learning Representations (ICLR), 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy, “Pick-a-Pic: An open dataset of user preferences for text-to-image generation,” arXiv, 2023, doi: 10.48550/arXiv.2305.01569.
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. R. Joty, and N. Naik, “Diffusion Model Alignment Using Direct Preference Optimization,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8228–8238, 2023.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv, 2022, doi: 10.48550/arXiv.2207.12598.
A. Ramesh, P. Dhariwal, A. Nichol, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents,” arXiv e-prints, 2022, doi: 10.48550/arXiv.2204.06125.
J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, and Y. Choi, “CLIPScore: A Reference-free Evaluation Metric for Image Captioning,” in EMNLP 2021, pp. 7514–7528, Association for Computational Linguistics, 2021, doi: 10.18653/v1/2021.emnlp-main.595.
S. He, Y. Zhang, R. Xie, D. Jiang, and A. Ming, “Rethinking Image Aesthetics Assessment: Models, Datasets and Benchmarks,” in International Joint Conference on Artificial Intelligence, 2022.
P. Liao, X. Li, X. Liu, and K. Keutzer, “The ArtBench Dataset: Benchmarking Generative Models with Artworks,” arXiv, abs/2206.11404, 2022.
S. Zhou, M. Ye, and W. Luo, “ISEM: Image Steganography Enhancement Model Based on Generative Adversarial Networks,” Chinese Journal of Network and Information Security, vol. 11, no. 5, pp. 126–136, 2025.
S. Yan, “Research on Immersive Cross-scenario Collaborative Innovative Design in Digital Intelligent Performance,” Journal of Beihang University (Social Sciences Edition), vol. 37, no. 5, pp. 51–59, 2024, doi: 10.13766/j.bhsk.1008-2204.2024.1171.
Y. Li and L. Zhang, “Research on Face Image Generation Method Based on CLIP Model and Text Reconstruction,” Journal of Test and Measurement Technology, vol. 38, no. 2, pp. 154–160, 2024.