Culturally-Guided Style Transfer for Southeast Asian Text using Constrained Diffusion Models
Main Article Content
Abstract
Preserving culturally grounded semantics during multilingual literary style transfer remains challenging because existing text generation models often fail to maintain narrative structure, symbolic imagery, and rhetorical consistency across heterogeneous linguistic contexts. This study proposes a culturally guided style transfer framework based on constrained diffusion models that jointly models semantic hierarchies and stylistic representations in a unified latent space. By introducing style-modifying latent variables into the diffusion process and coupling them with hierarchical semantic embeddings, the proposed approach enables synchronized regulation of cultural semantics and literary style throughout forward diffusion and reverse denoising. Unlike conventional conditional generation methods, stylistic evolution emerges through semantic-aware latent perturbations, improving the preservation of cultural identity while maintaining semantic coherence. Experiments on five Southeast Asian literary corpora demonstrate stable cultural mapping and consistent style reconstruction, achieving high levels of cultural image consistency together with reliable narrative and rhetorical preservation across multilingual settings. Beyond literary generation, the proposed semantic-driven diffusion framework offers a transferable paradigm for intelligent information representation and cross-modal content encoding, providing potential insights for semantic communication, electromagnetic information transmission, and AI-assisted cultural computing where robust semantic preservation is essential.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
Q. Yi, X. Chen, C. Zhang, Z. Zhou, L. Zhu, and X. Kong, “Diffusion models in text generation: a survey,” PeerJ Computer Science, vol. 10, no. 1, pp. e1905, 2024, doi: 10.7717/peerj-cs.1905.
S. Colella, “The Language of the Digital Air”: AI-Generated Literature and the Performance of Authorship,” Humanities, vol. 14, no. 8, pp. 164, 2025, doi: 10.3390/h14080164.
S. Pang, X. Chen, Y. Xie, H. Zhan, B. Yin, and Y. Lu, “Diff-TST: Diffusion model for one-shot text-image style transfer,” Expert Systems with Applications, vol. 263, no. 1, pp. 125747, 2025, doi: 10.1016/j.eswa.2024.125747.
Y. Wu, H. Zhao, W. Chen, Y. Yang, and J. Bu, “TextStyler: A CLIP-based approach to text-guided style transfer,” Computers & Graphics, vol. 119, no. 1, pp. 103887, 2024, doi: 10.1016/j.cag.2024.103887.
E. Troiano, A. Velutharambath, and R. Klinger, “From theories on styles to their transfer in text: Bridging the gap with a hierarchical survey,” Natural Language Engineering, vol. 29, no. 4, pp. 849-908, 2023, doi: 10.1017/S1351324922000407.
J. Peng, Y. Zhou, X. Sun, L. Cao, Y. Wu, and F. Huang, “Knowledge-driven generative adversarial network for text-to-image synthesis,” IEEE Transactions on Multimedia, vol. 24, no. 1, pp. 4356-4366, 2021, doi: 10.1109/TMM.2021.3116416.
Q. Wang, S. Li, Z. Wang, X. Zhang, and G. Feng, “Multi-source style transfer via style disentanglement network,” IEEE Transactions on Multimedia, vol. 26, no. 1, pp. 1373-1383, 2023, doi: 10.1109/TMM.2023.3281087.
N. Ichien, D. Stamenković, and J. Holyoak K, “Large language model displays emergent ability to interpret novel literary metaphors,” Metaphor and Symbol, vol. 39, no. 4, pp. 296-309, 2024, doi: 10.1080/10926488.2024.2380348.
T. Chay-intr, H. Kamigaito, and M. Okumura, “Character-based Thai word segmentation with multiple attentions,” Journal of Natural Language Processing, vol. 30, no. 2, pp. 372-400, 2023, doi: 10.5715/jnlp.30.372.
A. Scalercio and A. Paes, “Masked transformer through knowledge distillation for unsupervised text style transfer,” Natural Language Engineering, vol. 30, no. 5, pp. 973-1008, 2024, doi: 10.1017/S1351324923000323.
Y. Shi, S. Zhang, C. Zhou, X. Liang, X. Yang, and L. Lin, “Gtae: graph transformer–based auto-encoders for linguisticconstrained text style transfer,” ACM Transactions on Intelligent Systems and Technology (TIST), vol. 12, no. 3, pp. 1-16, 2021, doi: 10.1145/3448733.
D. Wu and C. Monz, “Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine Translation,” arXiv e-prints, vol. 2305, no. 1, pp. 14189, 2023, doi: 10.18653/v1/2023.emnlp-main.605.
S. Cao, J. Li, P. Nelson K, and A. Kon M, “Coupled VAE: Improved accuracy and robustness of a variational autoencoder,” Entropy, vol. 24, no. 3, pp. 423, 2022, doi: 10.3390/e24030423.
H. Zhang, M. Wang, L. Zhang, Y. Wu, and Y. Luo, “Image Generation with Global Photographic Aesthetic Based on Disentangled Generative Adversarial Network,” Applied Sciences, vol. 13, no. 23, pp. 12871, 2023, doi: 10.3390/app132312871.
X. Han, Y. Wu, and R. Wan, “A method for style transfer from artistic images based on depth extraction generative adversarial network,” Applied Sciences, vol. 13, no. 2, pp. 867, 2023, doi: 10.3390/app13020867.
W. Lin J, W. Su T, and C. Chang C, “Chinese Story Generation Based on Style Control of Transformer Model and Content Evaluation Method,” Algorithms, vol. 18, no. 3, pp. 168, 2025, doi: 10.3390/a18030168.
R. Zandie and H. Mahoor M, “Topical language generation using transformers,” Natural Language Engineering, vol. 29, no. 2, pp. 337-359, 2023, doi: 10.1017/S1351324922000031.
H. Li, F. Xu, and Z. Lin, “ET-DM: Text to image via diffusion model with efficient Transformer,” Displays, vol. 80, no. 1, pp. 102568-102577, 2023, doi: 10.1016/j.displa.2023.102568.
C. Li, L. Zhang, and Q. Zheng, “Utilizing latent diffusion model to accelerate sampling speed and enhance text generation quality,” Electronics, vol. 13, no. 6, pp. 1093, 2024, doi: 10.3390/electronics13061093.
A. Wahida and M. H. Himawan, “Contemporary textile design creation sourced from visual aesthetics of Kawung motif classical batik,” Mudra Jurnal Seni Budaya, vol. 39, no. 4, pp. 494-507, 2024, doi: 10.31091/mudra.v39i4.1807.
H. An M and A. R. Jang, “Development of textile pattern design by MC Escher’s tessellation technique using chaekgeori icons,” Fashion and Textiles, vol. 10, no. 1, pp. 15, 2023, doi: 10.1186/s40691-023-00336-w.
W. Ke, Y. Guo, Q. Liu, W. Chen, P. Wang, H. Luo, et al., “MDM: Meta diffusion model for hard-constrained text generation,” Knowledge-Based Systems, vol. 283, no. 1, pp. 111147, 2024, doi: 10.1016/j.knosys.2023.111147.
C. Liu, W. Yin, Y. Xu, Q. Zhan, D. Zhang, and R. Ahmad, “DDM4TST: Diffusion Model for Fine-grained Text Style Transfer by Disentangled Representation,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 24, no. 10, pp. 1-17, 2025, doi: 10.1145/3749195.
Y. Wang, B. Zhang, W. Liu, J. Cai, and H. Zhang, “STMAP: A novel semantic text matching model augmented with embedding perturbations,” Information Processing & Management, vol. 61, no. 1, pp. 103576, 2024, doi: 10.1016/j.ipm.2023.103576.
J. Huang, D. Zheng, J. Li, and Patrick Cheong-Iao Pang, “EmbDiffounder: Semantic-Transfer Enhanced Sequence Diffusion Model for Text Generation,” IEEE Access, vol. 13, no. 1, pp. 205521-205531, 2025, doi: 10.1109/ACCESS.2025.3638810.
A. Asperti, D. Evangelista, S. Marro, and F. Merizzi, “Image embedding for denoising generative models,” Artificial Intelligence Review, vol. 56, no. 12, pp. 14511-14533, 2023, doi: 10.1007/s10462-023-10504-5.
Z. Yu, J. Jin, J. Zhao, Z. Fu, and J. Yang, “TtfDiffusion: Training-free and text-free image editing in diffusion models with structural and semantic disentanglement,” Neurocomputing, vol. 619, no. 1, pp. 129159, 2025, doi: 10.1016/j.neucom.2024.129159.
S. Chen, H. Zhang, M. Guo, Y. Lu, P. Wang, and Q. Qu, “Exploring low-dimensional subspace in diffusion models for controllable image editing,” Advances in neural information processing systems, vol. 37, no. 1, pp. 27340-27371, 2024, doi: 10.52202/079017-0859.
N. Gao, Y. Lu, P. Chen, G. Sun, R. Liang, and Y. Zhang, “TWDT: Training-Free Word-Level Controllable Diffusion Model for Text Generation,” Knowledge-Based Systems, vol. 330, no. 1, pp. 114437, 2025, doi: 10.1016/j.knosys.2025.114437.
Y. Xu, M. Gong, S. Xie, W. Wei, M. Grundmann, and K. Batmanghelich, “Semi-implicit denoising diffusion models (siddms),” Advances in neural information processing systems, vol. 36, no. 1, pp. 17383, 2024, doi: 10.48550/arxiv.2306.12511.
W. Han, H. Cho, D. Kim, and J. Y. Kim, “SAL-PIM: A Subarray-level processing-in-memory architecture with LUT-based linear interpolation for transformer-based text generation,” IEEE Transactions on Computers, vol. 74, no. 9, pp. 2909-2922, 2025, doi: 10.1109/TC.2025.3576935.
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Berg, “Structured denoising diffusion models in discrete state-spaces,” Advances in neural information processing systems, vol. 34, no. 1, pp. 17981-17993, 2021, doi: 10.48550/arXiv.2107.03006.
R. Jiang, C. Zheng G, T. Li, R. Yang T, D. Wang J, and X. Li, “A survey of multimodal controllable diffusion models,” Journal of Computer Science and Technology, vol. 39, no. 3, pp. 509-541, 2024, doi: 10.1007/s11390-024-3814-0.
C. Scribano, D. Pezzi, G. Franchini, and M. Prato, “Denoising diffusion models on model-based latent space,” Algorithms, vol. 16, no. 11, pp. 501, 2023, doi: 10.3390/a16110501.
T. Choi H, K. Nakamura, and W. Hong B, “Decoupled Latent Diffusion Model for Enhancing Image Generation,” IEEE Access, vol. 13, no. 1, pp. 130505-130516, 2025, doi: 10.1109/ACCESS.2025.3592163.