Structured Creative Evaluation for Text-to-Image Generative AI Models
Main Article Content
Abstract
In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practical for real deployment. To address this issue, this paper proposes a multi-dimensional image quality assessment framework for T2I tasks. The framework examines generated images from five dimensions—text fidelity, perceptual quality, object consistency, relational consistency, and global semantic alignment—and derives a final quality score through normalization and weighted fusion. In terms of methodology, the framework combines Tesseract OCR, perceptual quality analysis based on Laplacian variance and exposure statistics, YOLO object detection, BLIP-based visual question answering, and CLIP image-text similarity, thereby forming a modular evaluation pipeline with diagnostic capability. Experiments on multiple mainstream T2I models and representative prompts show that the proposed method can not only distinguish overall performance differences across models, but also provide interpretable results at the level of individual dimensions.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
I. Goodfellow, J. Pouget-Abadie, M. Mirza, et al., “Generative adversarial nets,” in Proc. Adv. Neural Inf. Process. Syst. (NIPS), vol. 27, 2014.
R. Rombach, A. Blattmann, D. Lorenz, et al., “High-resolution image synthesis with latent diffusion models,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 10684–10695, doi: 10.1109/CV PR52688.2022.01042.
A. Ramesh, M. Pavlov, G. Goh, et al., “Zero-shot text-to-image generation,” in Proc. Int. Conf. Mach. Learn. (ICML), 2021, pp. 8821–8831, doi: 10.48550/arXiv.2102.12092.
C. Saharia, W. Chan, S. Saxena, et al., “Photorealistic text-to-image diffusion models with deep language understanding,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2022, doi: 10.48550/arXiv.2205.11487.
Z. Wang, A. C. Bovik, H. R. Sheikh, et al., “Image quality assessment: from error visibility to structural similarity,” IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, Apr. 2004, doi: 10.1109/TIP.2003.819861.
T. Salimans, I. Goodfellow, W. Zaremba, et al., “Improved techniques for training GANs,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2016, doi: 10.48550/arXiv.1606.03498.
M. Heusel, H. Ramsauer, T. Unterthiner, et al., “GANs trained by a two time-scale update rule converge to a local Nash equilibrium,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2017, doi: 10.48550/arXiv.1706.08500.
A. Radford, J. W. Kim, C. Hallacy, et al., “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn. (ICML), 2021, pp. 8748–8763, doi: 10.48550/arXiv.2103.00020.
J. Li, D. Li, C. Xiong, et al., “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in Proc. Int. Conf. Mach. Learn. (ICML), 2022, pp. 12888–12900, doi: 10.48550/arXiv.2201.12086.
J. Li, D. Li, S. Savarese, et al., “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in Proc. Int. Conf. Mach. Learn. (ICML), 2023, pp. 19730–19742, doi: 10.48550/arXiv.2301.12597.
Y. Hu, Z. Luo, D. Zhang, et al., “TIFA: Accurate and interpretable text-to-image faithfulness evaluation with question answering,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 20413–20422, doi: 10.48550/arXiv.2307.06350.
K. Huang, K. Sun, E. Xie, et al., “T2I-CompBench: A comprehensive benchmark for open-world compositional text-to-image generation,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2023, doi: 10.48550/arXiv.2307.06350.
D. Ghosh, A. Gupta, and Y. Bengio, “GenEval: An object-focused framework for evaluating text-to-image alignment,” in Adv. Neural Inf. Process. Syst. Datasets Benchmarks Track (NeurIPS), 2023, doi: 10.48550/arXiv.2
J. Xu, X. Liu, Y. Wu, et al., “ImageReward: Learning and evaluating human preferences for text-to-image generation,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2023, doi: 10.48550/arXiv.2304.05977.
Y. Kirstain, A. Polyak, U. Singer, et al., “Pick-a-Pic: An open dataset of user preferences for text-to-image generation,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2023, doi: 10.48550/arXiv.2305.01569.
H. Talebi and P. Milanfar, “NIMA: Neural image assessment,” IEEE Trans. Image Process., vol. 27, no. 8, pp. 3998–4011, Aug. 2018, doi: 10.1109/TIP.2018.2831899.
J. Ke, Q. Wang, Y. Wang, et al., “MUSIQ: Multi-scale image quality transformer,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021, pp. 5148–5157, doi: 10.1109/ICCV48922.2021.00510.
J. Hessel, A. Holtzman, M. Forbes, et al., “CLIPScore: A reference-free evaluation metric for image captioning,” in Proc. Conf. Empir. Methods Nat. Lang. Process. (EMNLP), 2021, pp. 7514–7528, doi: 10.18653/v1/20 21.emnlp-main.595.
Q. Huynh-Thu and M. Ghanbari, “Scope of validity of PSNR in image/video quality assessment,” Electron. Lett., vol. 44, no. 13, pp. 800– 801, Jun. 2008, doi: 10.1049/el:20080522.
J. Redmon, S. Divvala, R. Girshick, et al., “You only look once: Unified, real-time object detection,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2016, pp. 779–788, doi: 10.1109/CVPR.2016.91.