A Thai Spoken Dialogue Generation System Incorporating Generative Adversarial Networks
Main Article Content
Abstract
Traditional Thai dialogue generation systems, under maximum likelihood training, often favor high-frequency "safe" sentence structures, resulting in monotonous style and poor multi-turn contextual coherence. Similar problems exist in automated quality control within the textile industry: traditional models, trained on large amounts of defect-free data, struggle to distinguish subtle or rare defects, and are difficult to capture long-range dependencies in multi-sensor data streams. This paper proposes a Thai spoken dialogue generation system based on RelGAN and Transformer. It captures multi-turn context through a segmented relative position self-attention encoder and segmented recurrent memory, while the decoder embeds a relational memory module to track dialogue states and long-term entity slots, simultaneously generating text and prosodic markers. The generated results are prosodic mapped using FastSpeech2 and synthesized into high-fidelity speech using HiFi-GAN. Experiments show that the system achieves a Distinct-2 score of 0.11, an intent retention rate of 97.5%, an average response time of 98.44 ms, and a MOS of 4.05. This method has significant potential for cross-industry applications, enabling the construction of highly consistent real-time diagnostic systems in the textile industry, processing continuous multi-cycle sensor data, and achieving rapid fault identification and feedback.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
T. Chen, “Design and Application of English Oral Online Dialogue System Based on Reinforcement Learning Algorithm,” Procedia Computer Science, vol. 261, no. 1, pp. 716–723, 2025, doi: https://doi.org/10.1016/j.procs.2025.04.325.
S. Bibauw, W. Van den Noortgate, T. Francois, and P. Desmet, “Dialogue systems for language learning: A meta-analysis,” Language Learning & Technology, vol. 26, no. 1, pp. 1–24, 2022, doi: https://doi.org/10.64152/10125/73488.
M. H. Hsu, P. S. Chen, and C. S. Yu, “Proposing a task-oriented chatbot system for EFL learners speaking practice,” Interactive Learning Environments, vol. 31, no. 7, pp. 4297–4308, 2023, doi: https://doi.org/10.1080/10494820.2021.1960864.
C. Zhai and S. Wibowo, “A WGAN-based dialogue system for embedding humor, empathy, and cultural aspects in education,” IEEE Access, vol. 11, no. 1, pp. 71940–71952, 2023, doi: https://doi.org/10.1109/ACCESS.2023.3294966.
A. C. Le and V. N. Huynh, “Enhancing conversational model with deep reinforcement learning and adversarial learning,” IEEE Access, vol. 11, no. 1, pp. 75955–75970, 2023, doi: https://doi.org/10.1109/ACCESS.2023.3297652.
M. Firdaus, N. Thangavelu, A. Ekbal, and P. Bhattacharyya, “I enjoy writing and playing, do you?: A personalized and emotion grounded dialogue agent using generative adversarial network,” IEEE Transactions on Affective Computing, vol. 14, no. 3, pp. 2127–2138, 2022, doi: https://doi.org/10.1109/TAFFC.2022.3155105.
M. Firdaus, A. Madasu, and A. Ekbal, “A unified framework for slot based response generation in a multimodal dialogue system,” Multimedia Tools and Applications, vol. 83, no. 4, pp. 11643–11667, 2024, doi: https://doi.org/10.1007/s11042-023-15915-8.
A. Fuad and M. Al-Yahya, “AraConv: Developing an Arabic task-oriented dialogue system using multi-lingual transformer model mT5,” Applied Sciences, vol. 12, no. 4, pp. 1881–1896, 2022, doi: https://doi.org/10.3390/app12041881.
X. Li, T. Liu, L. Zhang, F. Alqahtani, and A. Tolba, “A transformer-BERT integrated model-based automatic conversation method under English context,” IEEE Access, vol. 12, no. 1, pp. 55757–55767, 2024, doi: https://doi.org/10.1109/ACCESS.2024.3388100.
Y. Zhao, B. Cheng, Y. Huang, and Z. Wan, “FluGCF: a fluent dialogue generation model with coherent concept entity flow,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, no. 1, pp. 853–867, 2023, doi: https://doi.org/10.1109/TASLP.2023.3340610.
K. Manohar and R. Rajan, “Improving speech recognition systems for the morphologically complex Malayalam language using subword tokens for language modeling,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2023, no. 1, pp. 47–71, 2023, doi: https://doi.org/10.1186/s13636-023-00313-7.
M. Bhadri, “BiLSTMs and BPE for English to Telugu CLIR,” J. Electrical Systems, vol. 20, no. 3, pp. 2022–2029, 2024, doi: https://doi.org/10.52783/jes.1798.
Y. Qiu, L. Dong, W. Zhang, H. Xing, and J. Huang, “A diffusion enhanced CRF and BiLSTM framework for accurate entity recognition,” Scientific Reports, vol. 15, no. 1, pp. 1–25, 2025, doi: https://doi.org/10.1038/s41598-025-04036-x.
Z. Sun and X. Li, “Named entity recognition model based on feature fusion,” Information, vol. 14, no. 2, pp. 133–145, 2023, doi: https://doi.org/10.3390/info14020133.
D. Mazitov, I. Alimova, and E. Tutubalina, “Named entity recognition in Russian using multi-task LSTM-CRF,” Journal of Mathematical Sciences, vol. 273, no. 4, pp. 595–604, 2023, doi: https://doi.org/10.1007/s10958-023-06521-y.
H. Zhang, H. Song, S. Li, M. Zhou, and D. Song, “A survey of controllable text generation using transformer-based pre-trained language models,” ACM Computing Surveys, vol. 56, no. 3, pp. 1–37, 2023, doi: https://doi.org/10.1145/3617680.
M. Thakkar and N. N. Pise, “Leveraging Transformer-based Pretrained Language model for Task-oriented dialogue system,” International Journal of Computers, vol. 8, no. 1, pp. 1–4, 2023.
K. Nassiri and M. Akhloufi, “Transformer models used for text-based question answering systems,” Applied Intelligence, vol. 53, no. 9, pp. 10602–10635, 2023, doi: https://doi.org/10.1007/s10489-022-04052-8.
A. Rahali and M. A. Akhloufi, “End-to-end transformer-based models in textual-based NLP,” Ai, vol. 4, no. 1, pp. 54–110, 2023, doi: https://doi.org/10.3390/ai4010004.
S. Lafkiar and N. En Nahnahi, “An end-to-end transformer-based model for Arabic question generation,” Multimedia Tools and Applications, vol. 84, no. 20, pp. 22009–22023, 2025, doi: https://doi.org/10.1007/s11042-024-19958-3.
R. Bhat and R. Nanjundegowda, “A review on comparative analysis of generative adversarial networks’ architectures and applications,” Journal of Robotics and Control (JRC), vol. 6, no. 1, pp. 53–64, 2025, doi: https://doi.org/10.18196/jrc.v6i1.24160.
C. Shao, Z. Ma, M. Zhang, and Y. Feng, “Beyond mle: Convex learning for text generation,” Advances in Neural Information Processing Systems, vol. 36, no. 1, pp. 8913–8936, 2023, doi: https://doi.org/10.52202/075280-0391.
J. Xu, X. Liu, J. Yan, D. Cai, H. Li, and J. Li, “Learning to break the loop: Analyzing and mitigating repetitions for neural text generation,” Advances in Neural Information Processing Systems, vol. 35, no. 11, pp. 3082–3095, 2022, doi: https://doi.org/10.52202/068431-0223.
I. A. M. Huijben, W. Kool, M. B. Paulus, and R. J. G. Van Sloun, “A review of the gumbel-max trick and its extensions for discrete stochasticity in machine learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 1353–1371, 2022, doi: https://doi.org/10.1109/TPAMI.2022.3157042.
I. Alsmadi, N. Aljaafari, M. Nazzal, S. Alhamed, A. H. Sawalmeh, C. P. Vizcarra, et al., “Adversarial machine learning in text processing: a literature survey,” IEEE Access, vol. 10, no. 1, pp. 17043–17077, 2022, doi: https://doi.org/10.1109/ACCESS.2022.3146405.
M. Reyad, A. M. Sarhan, and M. Arafa, “A modified Adam algorithm for deep neural network optimization,” Neural Computing and Applications, vol. 35, no. 23, pp. 17095–17112, 2023, doi: https://doi.org/10.1007/s00521-023-08568-z.
H. Kabiri, Y. Ghanou, H. Khalifi, and G. Casalino, “AMAdam: adaptive modifier of Adam method,” Knowledge and Information Systems, vol. 66, no. 6, pp. 3427–3458, 2024, doi: https://doi.org/10.1007/s10115-023-02052-9.
C. Zhang, Y. Shao, H. Sun, L. Xing, Q. Zhao, and L. Zhang, “The WuC-Adam algorithm based on joint improvement of Warmup and cosine annealing algorithms,” Math. Biosci. Eng, vol. 21, no. 1, pp. 1270–1285, 2024, doi: https://doi.org/10.3934/mbe.2024054.