Colour Style Generation for Digital Images: A Controllable Diffusion Model with Semantic Segmentation Constraints
Main Article Content
Abstract
Artificial intelligence generation techniques are increasingly reshaping digital image colour editing by enabling flexible and controllable style reconstruction. A controllable diffusion model with semantic segmentation constraints is proposed to address semantic mismatch, boundary colour bleeding, and insufficient regional control in global colour grading. Built on a latent diffusion framework, the model integrates semantic segmentation, multi-scale regional feature encoding, Lab-space colour statistics, style condition embedding, cross-attention injection, semantic-mask-guided denoising, and multi-objective consistency optimisation. Region-specific colour constraints are dynamically introduced during reverse diffusion to preserve object structure, coordinate local tones, and support continuous adjustment of style intensity, hue, and saturation. Experiments show that the proposed model achieves an FID of 23.7, 19.4% lower than ControlNet, while SSIM reaches 0.891. Compared with ControlNet, LPIPS decreases from 0.171 to 0.142 and colour error ∆E drops from 8.2 to 6.8. Semantic mIoU increases to 82.6%, boundary F1 reaches 87.4%, and style similarity improves to 0.913, confirming superior generation quality, semantic preservation, and controllability. Ablation results verify complementary contributions from segmentation, embedding, consistency, and boundary smoothing.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
R. Rombach, A. Blattmann, D. Lorenz, et al., “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10684– 10695.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 3836–3847.
G. Couairon, J. Verbeek, H. Schwenk, et al., “Diffedit: Diffusion-based semantic image editing with mask guidance,” arXiv preprint arXiv:2210.11427, 2022.
G. Kwon and J. C. Ye, “Diffusion-based image translation using disentan-gled style and content representation,” arXiv preprint arXiv:2209.15264, 2022.
B. Kawar, S. Zada, O. Lang, et al., “Imagic: Text-based real image editing with diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 6007–6017.
T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 18392– 18402.
O. Avrahami, D. Lischinski, and O. Fried, “Blended diffusion for text-driven editing of natural images,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 18208–18218.
G. C. Tarrés, D. Ruta, T. Bui, et al., “Parasol: Parametric style control for diffusion image synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 2432–2442.
S. Huang, Q. Li, J. Liao, et al., “Controllable image synthesis methods, applications and challenges: a comprehensive survey,” Artificial Intelligence Review, vol. 57, no. 12, p. 336, 2024.
T. T. Nguyen-Quynh, S. H. Kim, and N. T. do, “Image colorization using the global scene-context style and pixel-wise semantic segmentation,” IEEE Access, vol. 8, pp. 214098–214114, 2020.
O. Tasar, S. L. Happy, Y. Tarabalka, et al., “ColorMapGAN: Unsupervised domain adaptation for semantic segmentation using color mapping generative adversarial networks,” IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 10, pp. 7178–7193, 2020.
Y. Yu, D. Li, B. Li, et al., “Multi-style image generation based on semantic image,” The Visual Computer, vol. 40, no. 5, pp. 3411–3426, 2024.
L. Qi, L. Yang, W. Guo, et al., “Unigs: Unified representation for image generation and segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6305–6315.
W. Luo, S. Yang, X. Zhang, et al., “Siedob: Semantic image editing by disentangling object and background,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1868– 1878.
D. Pavllo, A. Lucchi, and T. Hofmann, “Controlling style and semantics in weakly-supervised image generation,” in European conference on computer vision. Cham: Springer International Publishing, 2020, pp. 482–499.
Y. Jia, L. Hoyer, S. Huang, et al., “Dginstyle: Domain-generalizable semantic segmentation with image diffusion models and stylized semantic control,” in European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024, pp. 91–109.
Z. Lin, Z. Wang, H. Chen, et al., “Image style transfer algorithm based on semantic segmentation,” IEEE Access, vol. 9, pp. 54518–54529, 2021.
Y. Zhao, Z. Zhong, N. Zhao, et al., “Style-hallucinated dual consistency learning for domain generalized semantic segmentation,” in European conference on computer vision. Cham: Springer Nature Switzerland, 2022, pp. 535–552.
X. Gong, S. Chen, B. Zhang, et al., “Style consistent image generation for nuclei instance segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 3994–4003.
Y. Yu, C. Wang, Q. Fu, et al., “Techniques and challenges of image segmentation: A review,” Electronics, vol. 12, no. 5, p. 1199, 2023.
A. Satchidanandam, R. M. S. Al Ansari, A. L. Sreenivasulu, et al., “Enhancing style transfer with GANs: perceptual loss and semantic segmentation,” Int. J. Adv. Comput. Sci. Appl, vol. 11, pp. 321–329, 2023.
T. Chen, Y. Yao, X. Huang, et al., “Spatial structure constraints for weakly supervised semantic segmentation,” IEEE Transactions on Image Processing, vol. 33, pp. 1136–1148, 2024.
K. Psychogyios, H. C. Leligou, F. Melissari, et al., “Samstyler: Enhancing visual creativity with neural style transfer and segment anything model (sam),” IEEe Access, vol. 11, pp. 100256–100267, 2023.