Reconstruction of Semantic Consistency in Brand Visual Identity System via GLIGEN-Guided Multimodal Diffusion Modeling
Main Article Content
Abstract
Addressing the fragmentation of the Brand Visual Identity System (VIS) in digital communication, especially for diverse textures, layouts and brand elements in digital communication, this paper proposes a GLIGEN-guided multimodal diffusion modeling reconstruction framework to achieve precise matching between brand semantics and visual presentation. By constructing a dual-layer label system of "core semantics - visual semantics", designing a five-layer architecture including input layer, feature fusion layer, GLIGEN guidance layer, diffusion generation layer and semantic verification layer, a quantitative evaluation system combining CLIP semantic similarity and expert research is established. Taking 10 brands (5 fast-moving consumer goods and 5 high-end service brands) as experimental objects, comparing with traditional design methods and diffusion models without GLIGEN guidance, the results show that the CLIP semantic similarity of the proposed model reaches 0.89±0.03, the SSIM visual consistency reaches 0.86±0.02, and the single-image generation time is only 8.2 seconds, which significantly improves the efficiency compared with traditional manual design. Case verification shows that this framework can increase the cognitive degree of fast-moving consumer goods brands by 23% and the cross-scenario visual unity of high-end service brands reaches 92%. This research fills the gap of multimodal generation technology in the field of brand semantic control and provides a technical paradigm for brand VIS reconstruction, enhancing the ability of brands to maintain aesthetic coherence across diverse digital presentations.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
G. Guo, “On the Translation Mechanism of Brand Visual Symbols,” Green Packaging, no. 09, pp. 117-120, 2025, doi: 10.19362/j.cnki.cn10-1400/tb.2025.09.023.
X. Liu, “Research on Brand Image Design and Dynamic Communication Strategy in the Digital Age,” Portrait Photography, no. 11, pp. 208-209, 2025.
Y. Shi and Peng Ting, “Research on Enterprise Visual Identity System (VIS) and Brand Building – Taking Lingang Group VIS Upgrade Project as an Example,” Enterprise Research, no. 04, pp. 20-23, 2025.
Y. Zhang, “Research on the Design and Communication Strategy of Brand Visual Identity System,” Gallery, no. 10, pp. 79-81, 2024.
Y. Li, H. Liu, Q. Wu, F. Mu, J. Yang, J. Gao, et al., “Gligen: Open-set grounded text-to-image generation,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition; 2023, pp. 22511-22521, [Online]. Available: https://gligen.github.io/.
Z. Jiang, L. M. Po, X. Xu, Y. Wang, H. Wu, Y. Liu, et al., “RMP-adapter: A region-based Multiple Prompt Adapter for multi-concept customization in text-to-image diffusion model,” Expert Systems with Applications, vol. 274, Art. no. 126936, 2025, doi: 10.1016/j.eswa.2025.126936.
D. A. Aaker, “Building strong brands,” New York: Simon and schuster; 2012, doi: 10.1057/palgrave.bm.2540156.
T. Chen, “The Impact of the Matching between Brand Logo Design and Brand Personality on Brand Equity,” Wuhan, China: Wuhan University Press; 2023.
Y. Jun and H. Lee, “Multisensory congruence in brand identity: evidence from the global automotive market,” Archives of Design Research, vol. 33, no. 4, pp. 19-40, 2020, doi: 10.15187/adr.2020.11.33.4.19.
X. Luo and B. Zhao, “2024 Generative AI Image Model Annual Report,” Art Studies, no. 01, pp. 145-156, 2025.
K. Zheng and D. Wang, “Application of Artificial Intelligence in the Field of Image Generation – Taking Stable Diffusion and ERNIE-ViLG as Examples,” Science & Technology Vision, no. 35, pp. 50-54, 2022.
H. Zhang, W. Yin, Y. Fang, L. Li, B. Duan, Z. Wu, et al., “Ernie-vilg: Unified generative pre-training for bidirectional vision-language generation,” arXiv preprint arXiv:2112.15283, 2021, doi: 10.48550/arXiv.2112.15283.
Y. Ma, “Research on the Application of Color Elements in the Construction of Brand Visual Symbols of Nijiazhuang Clay Sculpture,” Color, no. 08, pp. 19-21, 2025.
M. Bai, “Research on the Application of Traditional Cultural Elements in Modern Brand Design,” Tiangong, no. 32, pp. 73-76, 2025.
Q. Liu, “Application of Brand Visual Identity System in Advertising Design,” Shanghai Fashion, no. 06, pp. 180-182, 2024.
M. Zhang, H. Sun, and H. Zhang, “Research on the Design of E-commerce Brand Visual Symbols from the Perspective of Semiotics,” Industrial Design Research, no. 00, pp. 309-324, 2022.
F. Su, “Analysis of the Application of Intangible Cultural Heritage Visual Elements in Brand Logo Design,” Shoe Craft and Design, vol. 5, no. 16, pp. 30-32, 2025.