Artistic Creation Inspiration Generation System Based on Diffusion Model
Main Article Content
Abstract
In traditional artistic creation, inspiration relies heavily on subjective experience and occasional associations, resulting in limited systematicity and controllability during the early conception stage. Similar challenges also exist in intelligent visual design and information presentation for engineering applications, where efficient human–machine interaction and semantic visualization are increasingly important for electromagnetic communication systems and digital design environments. This paper proposes a text-guided diffusion model for artistic inspiration generation, particularly suitable for visual creation tasks such as illustration and conceptual art while remaining applicable to product innovation fields. The proposed framework integrates CLIP semantic embedding with conditional control mechanisms into a customized UNet architecture based on Stable Diffusion, enabling user text prompts to guide image synthesis through multi-layer cross-attention. Classifier-free guidance (CFG) and a multi-scale noise scheduling strategy are further introduced to improve style controllability, detail representation, and generation stability while supporting iterative refinement. Experimental results demonstrate stable performance with high semantic consistency (CLIP similarity of 0.82), effective style control (0.76), and favorable image quality (FID of 18.5). User evaluations achieve scores of 4.33 for creative inspiration and 4.49 for overall satisfaction, confirming that the proposed system provides an efficient and expressive visual inspiration generation tool. The framework also offers a systematic approach for multimodal semantic visualization and intelligent content generation, which may provide useful references for visual information representation and human–machine collaborative design in advanced electromagnetic and communication-related applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
V. El Ardeliya, J. Taylor, and J. Wolfson, “Exploration of artificial intelligence in creative fields: Generative art, music, and design,” International Journal of Cyber and IT Service Management, vol. 4, no. 1, pp. 40-46, 2024, doi: 10.34306/ijcitsm.v4i1.149.
Y. Li and Q. Zhang, “The Analysis of Aesthetic Preferences for Cultural and Creative Design Trends under Artificial Intelligence,” IEEE Access, vol. 12, pp. 158799-158808, 2024, doi: 10.1109/ACCESS.2024.3486031.
M. Monser and E. Fadel, “A modern vision in the applications of artificial intelligence in the field of visual arts,” International Journal of Multidisciplinary Studies in Art and Technology, vol. 6, no. 1, pp. 73-104, 2023, doi: 10.21608/ijmsat.2024.274900.1021.
A. Oksanen, A. Cvetkovic, N. Akin, R. Latikka, J. Bergdahl, Y. Chen, et al., “Artificial intelligence in fine arts: A systematic review of empirical research,” Computers in Human Behavior: Artificial Humans, vol. 1, no. 2, Art. no. 100004, 2023, doi: 10.1016/j.chbah.2023.100004.
B. Wang, Q. Chen, and Z. Wang, “Diffusion-based visual art creation: A survey and new perspectives,” ACM Computing Surveys, vol. 57, no. 10, pp. 1-37, 2025, doi: 10.1145/3728459.
N. Dehouche and K. Dehouche, “What’s in a text-to-image prompt? The potential of stable diffusion in visual arts education,” Heliyon, vol. 9, no. 6, Art. no. e16757-e16769, 2023, doi: 10.1016/j.heliyon.2023.e16757.
Y. Jiang, S. Yang, H. Qiu, W. Wu, C. C. Loy, and Z. Liu, “Text2human: Text-driven controllable human image generation,” ACM Transactions on Graphics (TOG), vol. 41, no. 4, pp. 1-11, 2022, doi: 10.1145/3528223.3530104.
R. Fridman, A. Abecasis, Y. Kasten, and T. Dekel, “Scenescape: Text-driven consistent scene generation,” in Oh A, Naumann T, Globerson A, Saenko K, Hardt M, Levine S. editor. Advances in Neural Information Processing Systems. Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’23), 10-16 December 2023, Red Hook, NY, USA. Curran Associates Inc, 2023, pp. 39897-39914, [Online]. Available: https://dl.acm.org/doi/10.52202/075280-1734.
R. Safta, “Artistic Inspiration and Divine Inspiration,” Limba S,i Literatura–Repere Identitare În Context European, vol. (32), pp. 175-186, 2023.
C. Chang, “Being inspired by media content: Psychological processes leading to inspiration,” Media Psychology, vol. 26, no. 1, pp. 72-87, 2023, doi: 10.1080/15213269.2022.2097927.
F. K. El-Shafey, E. L. Mohamed, A. M. Fouad, A. S. Elkhayat, and A. G. Hassabo, “The Influence of Nature on Art and Graphic Design: The Connection with Raw Materials and Prints,” Journal of Textiles, Coloration and Polymer Science, vol. 21, no. 2, pp. 385-396, 2024, doi: 10.21608/jtcps.2024.259042.1276.
D. Gamberini, ““Unjustly Tormented by Love”: Eros as a Source of Artistic Inspiration in an Epigram for Gian Giorgio Lascaris, Alias Pyrgoteles,” Source: Notes in the History of Art, vol. 41, no. 3, pp. 176-185, 2022, doi: 10.1086/720923.
F. Tao, “A new harmonisation of art and technology: Philosophic interpretations of artificial intelligence art,” Critical Arts, vol. 36, no. 1-2, pp. 110-125, 2022, doi: 10.1080/02560046.2022.2112725.
L. Long, C. Xinyi, N. Ruoyu W E, T. J. J. Li, and L. C. Ray, “Sketchar: supporting character design and illustration prototyping using generative AI,” Proceedings of the ACM on Human-Computer Interaction, vol. 8, no. CHI PLAY, pp. 1-28, 2024, doi: 10.1145/3677102.
J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans, “Cascaded diffusion models for high fidelity image generation,” Journal of Machine Learning Research, vol. 23, no. 47, pp. 1-33, 2022, doi: papers/v23/21-0635.html.
F. A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Transactions on Patern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10850-10869, 2023, doi: 10.1109/TPAMI.2023.3261988.
D. H. Kwon, “Analysis of prompt elements and use cases in image-generating AI: Focusing on Midjourney, Stable Diffusion, Firefly, DALL· E,” Journal of Digital Contents Society, vol. 25, no. 2, pp. 341-354, 2024, doi: 10.9728/dcs.2024.25.2.341.
Y. Hao, Z. Chi, L. Dong, and F. Wei, “Optimizing prompts for text-to-image generation,” in Oh A, Naumann T, Globerson A, Saenko K, Hardt M, Levine S. editor. Advances in Neural Information Processing Systems. Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’23), 10-16 December 2023, Red Hook, NY, USA. Curran Associates Inc, 2023, pp. 66923–66939, [Online]. Available: https://dl.acm.org/doi/10.52202/075280-2923.