Enhancing Creative Expression Skills in Digital Media Education Using CLIP-Driven Multimodal Generation Systems
Main Article Content
Abstract
Current digital media education lacks a quantitative assessment method for creativity that simultaneously considers semantic fidelity, aesthetic quality, and stylistic diversity. This paper proposes a CLIP-driven multimodal generation and evaluation framework for improving creative expression skills. In the generation phase, the CLIP text encoder extracts semantic vectors from student prompts and feeds them into a diffusion model as conditional inputs. During sampling, embeddings of intermediate images and corresponding prompts are compared to guide semantic correction. In the assessment phase, CLIPScore is used to quantify semantic fidelity, a neural image assessment model outputs aesthetic quality scores, and principal component analysis is performed on CLIP visual embeddings from multi-round generated images to compute style dispersion through the covariance matrix trace. These three indicators are weighted to form a comprehensive creative-expression score. Propensity score matching and difference-in-differences estimation are further applied to identify the net causal effect of the system intervention. Experiments involving digital media students show a median semantic fidelity of 0.631, mean visual expressiveness of 0.774, style exploration breadth of 0.629, and average comprehensive creative-expression score of 0.782. The framework enables three-dimensional quantitative assessment and causal evaluation of creative expression in digital media education.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
M. Liu, W. Pang, J. Guo, and Y. Zhang, “A meta-analysis of the effect of multimedia technology on creative performance,” Education and Information Technologies, vol. 27, no. 6, pp. 8603-8630, 2022, DOI: 10.1007/s10639-022-10981-1.
H. Long, A. Kerr B, E. Emler T, and M. Birdnow, “A critical review of assessments of creativity in education,” Review of Research in Education, vol. 46, no. 1, pp. 288-323, 2022, DOI: 10.3102/0091732X221084326.
Y. Zhang and Z. Li, “Assessing GenAI-assisted digital multimodal composing: Reconceptualizing a genre-based framework through self-assessment and peer assessment,” Assessing Writing, vol. 68, no. 1, pp. 101017-101031, 2026, DOI: 10.1016/j.asw.2026.101017.
C. Laupichler M, A. Aster, O. Perschewski J, and J. Schleiss, “Evaluating AI courses: A valid and reliable instrument for assessing artificial-intelligence learning through comparative self-assessment,” Education Sciences, vol. 13, no. 10, pp. 978-995, 2023, DOI: 10.3390/educsci13100978.
Y. Zhuang, R. Zhao, Z. Xie, and L. Yu P, “Enhancing language learning through generative AI feedback on picture-cued writing tasks,” Computers and Education: Artificial Intelligence, vol. 1, no. 1, pp. 100450-100462, 2025, DOI: 10.1016/j.caeai.2025.100450.
J. Jiang, “When generative artificial intelligence meets multimodal composition: Rethinking the composition process through an AI-assisted design project,” Computers and Composition, vol. 74, no. 1, pp. 102883-102896, 2024, DOI: 10.1016/j.compcom.2024.102883.
A. Ringvold T, I. Strand, P. Haakonsen, and S. Strand K, “The AI generative text-to-image creative learning process: An art and design educational perspective,” Design and Technology Education: An International Journal, vol. 29, no. 2, pp. 359-379, 2024, DOI: 10.24377/DTEIJ.article2433.
G. Donnici, G. Galiè, and L. Frizziero, “Rethinking Sketching: Integrating Hand Drawings, Digital Tools, and AI in Modern Design,” Designs, vol. 9, no. 5, pp. 119-131, 2025, DOI: 10.3390/designs9050119.
L. Li, T. Zhu, P. Chen, Y. Yang, Y. Li, and W. Lin, “Image aesthetics assessment with attribute-assisted multi-modal memory network,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 12, pp. 7413-7424, 2023, DOI: 10.1109/TCSVT.2023.3272984.
J. Liu, “Modeling analysis of perceived differences in product appearance design based on visual communication,” International Journal for Simulation and Multidisciplinary Design Optimization, vol. 16, no. 1, pp. 12-26, 2025, DOI: 10.1051/smdo/2025016.
Z. Tan, X. Yang, Z. Ye, Q. Wang, Y. Yan, A. Nguyen, and K. Huang, “Semantic similarity distance: Towards better text-image consistency metric in text-to-image generation,” Pattern Recognition, vol. 144, no. 1, pp. 109883-109896, 2023, DOI: 10.1016/j.patcog.2023.109883.
F. Jahara, M. Khoshnoodi, Y. Lu, M. Saxon, A. Sharma, and Y. Wang W, “Who evaluates the evaluations? objectively scoring text-to-image prompt coherence metrics with t2iscorescore (ts2),” Advances in Neural Information Processing Systems, vol. 37, no. 1, pp. 85630-85657, 2024, DOI: 10.52202/079017-2719.
R. Gal, O. Patashnik, H. Maron, H. Bermano A, G. Chechik, and D. Cohen-Or, “Stylegan-nada: Clip-guided domain adaptation of image generators,” ACM Transactions on Graphics (TOG), vol. 41, no. 4, pp. 1-13, 2022, DOI: 10.1145/3528223.3530164.
H. Ma, M. Li, J. Yang, O. Patashnik, D. Lischinski, D. Cohen-Or, and H. Huang, “CLIP-Flow: Decoding images encoded in CLIP space,” Computational Visual Media, vol. 10, no. 6, pp. 1157-1168, 2024, DOI: 10.1007/s41095-023-0375-z.
I. Santos, A. Casal M, J. Correia, Á. Torrente-Patiño, P. Machado, and J. Romero, “Towards robust evaluation of aesthetic and photographic quality metrics: Insights from a comprehensive dataset,” Complexity, vol. 2024, no. 1, Art. no. 8223586, 2024, DOI: 10.1155/2024/8223586.
M. Daryanavard Chounchenani, A. Shahbahrami, R. Hassanpour, and G. Gaydadjiev, “Deep learning based image aesthetic quality assessment-a review,” ACM Computing Surveys, vol. 57, no. 7, pp. 1-36, 2025, DOI: 10.1145/3716820.
J. Thomson T, D. Pfurtscheller, K. Lobinger, N. Laba, and K. Christ, “Visual literacies in transition: multimodal co-production with generative AI,” Journal of Visual Literacy, vol. 44, no. 4, pp. 384-405, 2025, DOI: 10.1080/1051144x.2025.2566592.
Y. Wang, “Self-Supervised CLIP-Based Image Recognition and Analysis for Electronic Data Forensics,” IEEE Access, vol. 14, no. 1, pp. 3062-3077, 2025, DOI: 10.1109/ACCESS.2025.3648849.
P. Shabari Nath and R. Chouhan, “Multimodal Quality Assessment of AI-Generated Images using BERT-CLIP Feature Fusion: S,” Nath, R. Chouhan. Signal, Image and Video Processing, vol. 19, no. 14, pp. 1254-1271, 2025, DOI: 10.1007/s11760-025-04830-0.
H. Su, Z. Li, and J. Hu, “AI-driven Art Learning System: Automated Color Assessment Through Image Recognition,” Computer-Aided Design and Applications, vol. 1, no. 1, pp. 147-161, 2025, DOI: 10.14733/cadaps.2026.147-161.
R. Tang, Q. Li, and S. Tang, “Comparison of visual features for image-based visibility detection,” Journal of Atmospheric and Oceanic Technology, vol. 39, no. 6, pp. 789-801, 2022, DOI: 10.1175/jtech-d-21-0170.1.
H. Ye, X. Ge, J. Wang, J. Fu, X. Xin, J. Xue, Y. Chen, et al., “Beyond efficient fine-tuning: Efficient hybrid fine-tuning of CLIP models guided by explainable ViT attention,” Information Processing & Management, vol. 63, no. 4, pp. 104628-104653, 2026, DOI: 10.1016/j.ipm.2026.104628.
L. Liu, L. Yang, F. Yang, F. Chen, and F. Xu, “CLIP-Driven Few-Shot Species-Recognition Method for Integrating Geographic Information,” Remote Sensing, vol. 16, no. 12, pp. 2238-2257, 2024, DOI: 10.3390/rs16122238.
K. Tyshchuk, P. Karpikova, A. Spiridonov, A. Prutianova, A. Razzhigaev, and A. Panchenko, “On isotropy of multimodal embeddings,” Information, vol. 14, no. 7, pp. 392-404, 2023, DOI: 10.3390/info14070392.
H. Zhao, G. Jin, X. Jiang, and M. Li, “SDE-RAE: CLIP-based realistic image reconstruction and editing network using stochastic differential diffusion,” Image and Vision Computing, vol. 139, no. 1, pp. 104836-104851, 2023, DOI: 10.1016/j.imavis.2023.104836.
J. Kim, E. Kim, and K. Lee, “GODiff: region-specific semantic editing with CLIP-guided diffusion models,” IEEE Access, vol. 13, no. 1, pp. 112818-112834, 2025, DOI: 10.1109/access.2025.3580263.
Y. Hu, W. Dong, Y. Zhang, and L. Lu, “Image aesthetic quality assessment: A method based on deep convolutional capsule network,” PLoS One, vol. 20, no. 9, Art. no. e0331897, 2025, DOI: 10.1371/journal.pone.0331897.
S. Yang, Z. Wang, G. Wang, Y. Ke, F. Qin, J. Guo, and L. Chen, “A self-supervised image aesthetic assessment combining masked image modeling and contrastive learning,” Journal of Visual Communication and Image Representation, vol. 101, no. 1, pp. 104184-104197, 2024, DOI: 10.1016/j.jvcir.2024.104184.
M. Alkanan and Y. Gulzar, “Enhanced corn seed disease classification: leveraging MobileNetV2 with feature augmentation and transfer learning,” Frontiers in Applied Mathematics and Statistics, vol. 9, no. 1, pp. 1320177-1320189, 2024, DOI: 10.3389/fams.2023.1320177.
K.S, “A, Singh R P, Panda M K, Palaniappan K,” An Ensemble Approach using Self-attention based MobileNetV2 for SAR classification. Procedia Computer Science, vol. 235, no. 1, pp. 3207-3216, 2024, DOI: 10.1016/j.procs.2024.04.303.
T. Fakoya J and O. Ajinaja M, “A Cross-Modal Approach to Enhancing Image Retrieval With Contrastive Language-Image Pretraining (CLIP)- Based Embeddings and Facebook AI Similarity Search (FAISS) Indexing,” Cureus Journals, vol. 3, no. 1, pp. 1-14, 2026, DOI: 10.7759/s44389-026-00049-3.
J. Olaniyan, F. Verkijika S, and C. Obagbuwa I, “Cross-Modal Few-Shot Learning via Siamese Similarity Networks on CLIP Embeddings for Fine-Grained Image Classification,” Applied Sciences, vol. 16, no. 7, pp. 3181-3192, 2026, DOI: 10.3390/app16073181.
G. Kwon and C. Ye J, “One-shot adaptation of gan in just one clip,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 10, pp. 12179-12191, 2023, DOI: 10.1109/tpami.2023.3283551.
Z. Zhou, S. Dong, C. Ding, X. Gao, Y. He, and Y. Gong, “Diversity covariance-aware prompt learning for vision-language models,” Pattern Recognition, vol. 1, no. 1, pp. 112806-112815, 2025, DOI: 10.1016/j.patcog.2025.112806.
I. Naimi A and W. Whitcomb B, “Defining and identifying average treatment effects,” American Journal of Epidemiology, vol. 192, no. 5, pp. 685-687, 2023, DOI: 10.1093/aje/kwad012.