Using DeepSeek-VL Model to Build Multi-Style Interior Decoration Design Generation and Semantic Consistency Enhancement Mechanisms
Main Article Content
Abstract
This paper proposes a multi-style interior decoration design generation framework based on SDXL, integrating the DeepSeek-VL model, a local cross-modal attention mechanism, and an example-driven semantic consistency enhancement mechanism to effectively address feature confusion and text-image semantic inconsistency during multistyle switching. Beyond intelligent interior visualization, the proposed framework provides reliable semantic modeling capabilities for digital environments that can support engineering-oriented applications such as electromagnetic-aware building planning and intelligent spatial design. Style vector embedding is employed for accurate style representation, while the SDXL conditional decoder enables multi-style feature fusion. DeepSeek-VL achieves global semantic alignment, the local cross-modal attention module verifies the spatial correspondence of key design elements, and the example-driven retrieval mechanism injects features from similar real samples into the UNet layer to improve detail generation. Experimental results demonstrate outstanding performance in semantic consistency (similarity 0.942), style discriminability (distribution entropy 0.88), and visual quality (average MOS 4.65). The proposed framework provides an effective solution for automated multi-style interior design while offering a reliable semantic generation paradigm for intelligent spatial modeling and engineering visualization in complex built environments.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
E. Joy and C. Raja, “Digital 3D modeling for preconstruction real-time visualization of home interior design through virtual reality,” Construction Innovation, vol. 24, no. 2, pp. 643-653, 2024, doi: 10.1108/CI-10-2020-0174.
H. Lee, S. Je, R. Kim, H. Verma, H. Alavi, and A. Bianchi, “Partitioning open-plan workspaces via augmented reality,” Personal and Ubiquitous Computing, vol. 26, no. 3, pp. 609-624, 2022, doi: 10.1007/s00779-019-01306-0.
J. H. Choi and J. S. Lee, “A study of interior style transformation with GAN model,” Journal of KIBIM, vol. 12, no. 1, pp. 55-61, 2022, doi: 10.13161/kibim.2022.12.1.055.
Z. Wu, X. Jia, R. Jiang, Y. Ye, H. Qi, and C. Xu, “CSID-GAN: A customized style interior floor plan design framework based on generative adversarial network,” IEEE Transactions on Consumer Electronics, vol. 70, no. 1, pp. 2353-2364, 2024, doi: 10.1109/TCE.2024.3376956.
J. Chen, Z. Shao, and B. Hu, “Generating interior design from text: A new diffusion model-based method for efficient creative design,” Buildings, vol. 13, no. 7, pp. 1861-1878, 2023, doi: 10.3390/buildings13071861.
J. Chen, Z. Shao, X. Zheng, K. Zhang, and J. Yin, “Integrating aesthetics and efficiency: AI-driven diffusion models for visually pleasing interior design generation,” Scientific Reports, vol. 14, no. 1, pp. 3496-3509, 2024, doi: 10.1038/s41598-024-53318-3.
J. Chen, X. Zheng, Z. Shao, M. Ruan, H. Li, D. Zheng, et al., “Creative interior design matching the indoor structure generated through diffusion model with an improved control network,” Frontiers of Architectural Research, vol. 14, no. 3, pp. 614-629, 2025, doi: 10.1016/j.foar.2024.08.003.
H. Cui, L. Liu, H. Zhang, D. Liu, Y. Ma, and Z. Wang, “Detail preserving image generation method based on semantic consistency,” Journal of Computer-Aided Design & Computer Graphics, vol. 34, no. 10, pp. 1497-1505, 2022, doi: 10.3724/SP.J.1089.2022.19724.
Y. Ma, L. Liu, H. Zhang, C. Wang, and Z. Wang, “Generative adversarial network based on semantic consistency for text-to-image generation,” Applied Intelligence, vol. 53, no. 4, pp. 4703-4716, 2023, doi: 10.1007/s10489-022-03660-8.
Y. Tewel, O. Kaduri, R. Gal, Y. Kasten, L. Wolf, G. Chechik, et al., “Training-free consistent text-to-image generation,” ACM Transactions on Graphics (TOG), vol. 43, no. 4, pp. 1-18, 2024, doi: 10.1145/3658157.
Z. Li and R. Zhang, “Direct Distillation: A Novel Approach for Efficient Diffusion Model Inference,” Journal of Imaging, vol. 11, no. 2, pp. 66-88, 2025, doi: 10.3390/jimaging11020066.
R. K. Tripti, P. Chethan, S. Arun Vikas, S. S. Monisha, and A. I. Deepseek Open-Source, “International Journal of Trend in Scientific Research and Development,” 2025; 9(3):555-560, [Online]. Available: http://eprints.umsida.ac.id/id/eprint/16117.
Z. Deng, W. Ma, Q. L. Han, W. Zhou, X. Zhu, S. Wen, et al., “Exploring DeepSeek: A Survey on Advances, Applications, Challenges and Future Directions,” IEEE/CAA Journal of Automatica Sinica, vol. 12, no. 5, pp. 872-893, 2025, doi: 10.1109/JAS.2025.125498.
J. Wu, J. Chen, J. Wu, W. Shi, X. Wang, and X. He, “Understanding contrastive learning via distributionally robust optimization,” Advances in Neural Information Processing Systems, vol. 36, no. 1, pp. 23297-23320, 2023, doi: 10.48550/arXiv.2310.11048.
S. Lee, J. Park, and J. Park, “CrossFormer: Cross-guided attention for multi-modal object detection,” Pattern Recognition Letters, vol. 179, no. 1, pp. 144-150, 2024, doi: 10.1016/j.patrec.2024.02.012.
F. Zheng, W. Li, X. Wang, L. Wang, X. Zhang, and H. Zhang, “A cross-attention mechanism based on regional-level semantic features of images for cross-modal text-image retrieval in remote sensing,” Applied Sciences, vol. 12, no. 23, pp. 1-19, 2022, doi: 10.3390/app122312221.
H. Tan, B. Yin, K. Xu, H. Wang, X. Liu, and X. Li, “Attention-bridged modal interaction for text-to-image generation,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 5400-5413, 2023, doi: 10.1109/TCSVT.2023.3347971.
F. Wang, Y. Su, R. Wang, J. Sun, F. Sun, and H. Li, “Cross-modal and cross-level attention interaction network for salient object detection,” IEEE Transactions on Artificial Intelligence, vol. 5, no. 6, pp. 2907-2920, 2023, doi: 10.1109/TAI.2023.3333827.
D. He and C. Xie, “Semantic image segmentation algorithm in a deep learning computer network,” Multimedia Systems, vol. 28, no. 6, pp. 2065-2077, 2022, doi: 10.1007/s00530-020-00678-1.
J. Wang, X. Zhang, T. Yan, and A. Tan, “Dpnet: Dual-pyramid semantic segmentation network based on improved deeplabv3 plus,” Electronics, vol. 12, no. 14, pp. 3161-3176, 2023, doi: 10.3390/electronics12143161.
D. Lin, Y. X. Peng, J. Meng, and W. S. Zheng, “Cross-modal adaptive dual association for text-to-image person retrieval,” IEEE Transactions on Multimedia, vol. 26, no. 1, pp. 6609-6620, 2024, doi: 10.1109/TMM.2024.3355644.
X. Zhao, Y. Tian, K. Huang, B. Zheng, and X. Zhou, “Towards efficient index construction and approximate nearest neighbor search in high-dimensional spaces,” Proceedings of the VLDB Endowment, vol. 16, no. 8, pp. 1979-1991, 2023, doi: 10.14778/3594512.3594527.
L. Li, J. Cai, and J. Xu, “A learned index for approximate kNN queries in high-dimensional spaces,” Knowledge and Information Systems, vol. 64, no. 12, pp. 3325-3342, 2022, doi: 10.1007/s10115-022-01742-0.
M. Reyad, A. M. Sarhan, and M. Arafa, “A modified Adam algorithm for deep neural network optimization,” Neural Computing and Applications, vol. 35, no. 23, pp. 17095-17112, 2023, doi: 10.1007/s00521-023-08568-z.
I. V. Modoranu, M. Safaryan, G. Malinovsky, E. Kurtic, T. Robert, P. Richtárik, et al., “Microadam: Accurate adaptive optimization with low space overhead and provable convergence,” Advances in Neural Information Processing Systems, vol. 37, no. 1, pp. 1-43, 2024, doi: 10.52202/079017-0001.
H. Kabiri, Y. Ghanou, H. Khalifi, and G. Casalino, “AMAdam: Adaptive modifier of Adam method,” Knowledge and Information Systems, vol. 66, no. 6, pp. 3427-3458, 2024, doi: 10.1007/s10115-023-02052-9.