Colour Style Generation for Digital Images: A Controllable Diffusion Model with Semantic Segmentation Constraints

Main Article Content

Y. Q. Chen
H. L. Du

Abstract

Artificial intelligence generation techniques are increasingly reshaping digital image colour editing by enabling flexible and controllable style reconstruction. A controllable diffusion model with semantic segmentation constraints is proposed to address semantic mismatch, boundary colour bleeding, and insufficient regional control in global colour grading. Built on a latent diffusion framework, the model integrates semantic segmentation, multi-scale regional feature encoding, Lab-space colour statistics, style condition embedding, cross-attention injection, semantic-mask-guided denoising, and multi-objective consistency optimisation. Region-specific colour constraints are dynamically introduced during reverse diffusion to preserve object structure, coordinate local tones, and support continuous adjustment of style intensity, hue, and saturation. Experiments show that the proposed model achieves an FID of 23.7, 19.4% lower than ControlNet, while SSIM reaches 0.891. Compared with ControlNet, LPIPS decreases from 0.171 to 0.142 and colour error ∆E drops from 8.2 to 6.8. Semantic mIoU increases to 82.6%, boundary F1 reaches 87.4%, and style similarity improves to 0.913, confirming superior generation quality, semantic preservation, and controllability. Ablation results verify complementary contributions from segmentation, embedding, consistency, and boundary smoothing.

Downloads

Download data is not yet available.

Article Details

How to Cite
Chen, Y. Q., & Du, H. L. (2026). Colour Style Generation for Digital Images: A Controllable Diffusion Model with Semantic Segmentation Constraints. Advanced Electromagnetics, 15(3), 10256–10263. https://doi.org/10.7716/aem.v15i3.4228
Section
Research Articles

References

R. Rombach, A. Blattmann, D. Lorenz, et al., “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10684– 10695.

L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 3836–3847.

G. Couairon, J. Verbeek, H. Schwenk, et al., “Diffedit: Diffusion-based semantic image editing with mask guidance,” arXiv preprint arXiv:2210.11427, 2022.

G. Kwon and J. C. Ye, “Diffusion-based image translation using disentan-gled style and content representation,” arXiv preprint arXiv:2209.15264, 2022.

B. Kawar, S. Zada, O. Lang, et al., “Imagic: Text-based real image editing with diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 6007–6017.

T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 18392– 18402.

O. Avrahami, D. Lischinski, and O. Fried, “Blended diffusion for text-driven editing of natural images,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 18208–18218.

G. C. Tarrés, D. Ruta, T. Bui, et al., “Parasol: Parametric style control for diffusion image synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 2432–2442.

S. Huang, Q. Li, J. Liao, et al., “Controllable image synthesis methods, applications and challenges: a comprehensive survey,” Artificial Intelligence Review, vol. 57, no. 12, p. 336, 2024.

T. T. Nguyen-Quynh, S. H. Kim, and N. T. do, “Image colorization using the global scene-context style and pixel-wise semantic segmentation,” IEEE Access, vol. 8, pp. 214098–214114, 2020.

O. Tasar, S. L. Happy, Y. Tarabalka, et al., “ColorMapGAN: Unsupervised domain adaptation for semantic segmentation using color mapping generative adversarial networks,” IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 10, pp. 7178–7193, 2020.

Y. Yu, D. Li, B. Li, et al., “Multi-style image generation based on semantic image,” The Visual Computer, vol. 40, no. 5, pp. 3411–3426, 2024.

L. Qi, L. Yang, W. Guo, et al., “Unigs: Unified representation for image generation and segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6305–6315.

W. Luo, S. Yang, X. Zhang, et al., “Siedob: Semantic image editing by disentangling object and background,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1868– 1878.

D. Pavllo, A. Lucchi, and T. Hofmann, “Controlling style and semantics in weakly-supervised image generation,” in European conference on computer vision. Cham: Springer International Publishing, 2020, pp. 482–499.

Y. Jia, L. Hoyer, S. Huang, et al., “Dginstyle: Domain-generalizable semantic segmentation with image diffusion models and stylized semantic control,” in European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024, pp. 91–109.

Z. Lin, Z. Wang, H. Chen, et al., “Image style transfer algorithm based on semantic segmentation,” IEEE Access, vol. 9, pp. 54518–54529, 2021.

Y. Zhao, Z. Zhong, N. Zhao, et al., “Style-hallucinated dual consistency learning for domain generalized semantic segmentation,” in European conference on computer vision. Cham: Springer Nature Switzerland, 2022, pp. 535–552.

X. Gong, S. Chen, B. Zhang, et al., “Style consistent image generation for nuclei instance segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 3994–4003.

Y. Yu, C. Wang, Q. Fu, et al., “Techniques and challenges of image segmentation: A review,” Electronics, vol. 12, no. 5, p. 1199, 2023.

A. Satchidanandam, R. M. S. Al Ansari, A. L. Sreenivasulu, et al., “Enhancing style transfer with GANs: perceptual loss and semantic segmentation,” Int. J. Adv. Comput. Sci. Appl, vol. 11, pp. 321–329, 2023.

T. Chen, Y. Yao, X. Huang, et al., “Spatial structure constraints for weakly supervised semantic segmentation,” IEEE Transactions on Image Processing, vol. 33, pp. 1136–1148, 2024.

K. Psychogyios, H. C. Leligou, F. Melissari, et al., “Samstyler: Enhancing visual creativity with neural style transfer and segment anything model (sam),” IEEe Access, vol. 11, pp. 100256–100267, 2023.