Intelligent Interpretation of Traditional Folk Tales through Vocalization: a Generative Model Based on Emotion-Driven and Lip-Syncing
Main Article Content
Abstract
This study presents an integrated multi-modal signal generation framework for intelligent vocalization of traditional folk tales, ensuring emotion-driven speech synthesis with precise lip synchronization. The proposed system models continuous narrative emotion trajectories and embeds them into phoneme-level acoustic generation, using differentiable temporal alignment to synchronize audio and visual outputs. Multi-channel signals—including acoustic features and lip motion vectors— are processed through recursive smoothing and alignment matrices to maintain temporal consistency across modalities. Experimental results demonstrate median phoneme-duration deviations of 19.4–22.8 ms and lip-sync errors within 0.031–0.034, indicating stable audio-visual synchronization under complex narrative rhythms. By interpreting the system as a multi-channel signal processing and closed-loop control framework, the method mirrors principles of precision-engineered communication and control systems, providing insights into temporal alignment, cross-modal consistency, and real-time feedback optimization. This approach offers an engineering-oriented perspective for the development of advanced audio-visual generation, multi-sensor signal integration, and temporally coherent generative models.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
K. Liu, J. Wei, and R. Yuan, “Text-guided dynamic mouth motion capturing for person-generic talking face generation,” Knowledge-Based Systems, vol. 316, no. 1, pp. 113354-113380, 2025, doi: 10.1016/j.knosys.2025.113354.
X. Wang, Y. Huo, and Y. Liu, “Multimodal Feature-Guided Audio-Driven Emotional Talking Face Generation,” Electronics, vol. 14, no. 13, pp. 2684-2704, 2025, doi: 10.3390/electronics14132684.
J. Wang, O. Zhang, and Y. Jiang, “Multimodal diffusion framework for collaborative text image audio generation and applications,” Scientific Reports, vol. 15, no. 1, pp. 20604-20618, 2025, doi: 10.1038/s41598-025-05794-4.
J. Xu, X. Wu, and Z. Zhang, “MARS: Multimodal-Assisted Refined Semantic Alignment,” Information Processing & Management, vol. 63, no. 1, 104292, 2026, doi: 10.1016/j.ipm.2025.104292.
D. Jiang, J. Chang, and L. You, “Audio-Driven Facial Animation with Deep Learning: A Survey,” Information, vol. 15, no. 11, pp. 675-698, 2024, doi: 10.3390/info15110675.
X. Ji, Z. Liao, and L. Dong, “3D facial animation driven by speech-video dual-modal signals,” Complex & Intelligent Systems, vol. 10, no. 5, pp. 5951-5964, 2024, doi: 10.1007/s40747-024-01481-5.
L. Yu, H. Xie, and Y. Zhang, “Multimodal learning for temporally coherent talking face generation with articulator synergy,” IEEE Transactions on Multimedia, vol. 24, no. 1, pp. 2950-2962, 2021.
J. Salas-Caceres, J. Lorenzo-Navarro, and D. Freire-Obregon, “Multimodal emotion recognition based on a fusion of audio visual information with temporal dynamics,” Multimedia tools and applications, vol. 84, no. 23, pp. 27327-27343, 2025, doi: 10.1007/s11042-024-20227-6.
S. Liang, R. Zhou, and Q. Yuan, “ECE-TTS: A Zero-Shot Emotion Text-to-Speech Model with Simplified and Precise Control,” Applied Sciences, vol. 15, no. 9, pp. 5108-5129, 2025, doi: 10.3390/app15095108.
S. Hwang J, H. Lee S, and W. Lee S, “HiddenSinger: High-quality singing voice synthesis via neural audio co-dec and latent diffusion models,” Neural Networks, vol. 181, no. 1, pp. 106762-106771, 2025, doi: 10.1016/j.neunet.2024.106762.
R. Wu, Y. Yu, and F. Zhan, “Audio-driven talking face generation with diverse yet realistic facial animations,” Pattern Recognition, vol. 144, no. 1, pp. 109865-109891, 2023, doi: 10.1016/j.patcog.2023.109865.
X. Luo, S. Takamichi, and Y. Saito, “Emotion-controllable Speech Synthesis Using Emotion Soft Label, Utterance-level Prosodic Factors, and Word-level Prominence,” APSIPA Transactions on Signal and Information Processing, vol. 13, no. 1, pp. 1-30, 2024, doi: 10.1561/116.00000242.
W. Zhao and Z. Yang, “An emotion speech synthesis method based on vits,” Applied Sciences, vol. 13, no. 4, pp. 2225-2236, 2023, doi: 10.3390/app13042225.
Y. Hono, K. Hashimoto, and K. Oura, “Sinsy: A deep neural network-based singing voice synthesis system,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, no. 1, pp. 2803-2815, 2021, doi: 10.1109/TASLP.2021.3104165.
Y. Zhao, M. Yang, and Y. Lin, “AI-enabled text-to-music generation: A comprehensive review of methods, frameworks, and future directions,” Electronics, vol. 14, no. 6, pp. 1197-1249, 2025, doi: 10.3390/electronics14061197.
L. Liu, J. Wang, and S. Chen, “VividWav2Lip: high-fidelity facial animation generation based on speech-driven lip synchronization,” Electronics, vol. 13, no. 18, pp. 3657-3675, 2024, doi: 10.3390/electronics13183657.
Y. Fan, Z. Lin, and J. Saito, “Joint audio-text model for expressive speech-driven 3d facial animation,” Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 5, no. 1, pp. 1-15, 2022, doi: 10.1145/3522615.
X. Liu, X. Liu, and P. Yang, “An approach to optimizing semantic consistency for text-to-digital human generation,” Engineering Applications of Artificial Intelligence, vol. 160, no. 1, 111909, 2025, doi: 10.1016/j.engappai.2025.111909.
G. Chen, X. Liang, and W. He, “Enhancing Human-Computer Interaction Through Decoupling Motion and Camera Control in Human-Centric Video Generation,” International Journal of Human-Computer Interaction, vol. 1, no. 1, pp. 1-15, 2025, doi: 10.1080/10447318.2025.2487877.
J. Yang, J. Liu, and K. Huang, “Single-and cross-lingual speech emotion recognition based on WavLM domain emotion embedding,” Electronics, vol. 13, no. 7, pp. 1380-1395, 2024, doi: 10.3390/electronics13071380.
T. Lodewyckx, F. Tuerlinckx, and P. Kuppens, “A hierarchical state space approach to affective dynamics,” Journal of mathematical psychology, vol. 55, no. 1, pp. 68-83, 2011, doi: 10.1016/j.jmp.2010.08.004.
S. Sadok, S. Leglaive, and L. Girin, “A multimodal dynamical variational autoencoder for audiovisual speech representation learning,” Neural Networks, vol. 172, no. 1, pp. 106120-106135, 2024, doi: 10.1016/j.neunet.2024.106120.
H. Barakat, O. Turk, and C. Demiroglu, “Deep learning-based expressive speech synthesis: a systematic review of approaches, challenges, and resources,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, pp. 11-44, 2024, doi: 10.1186/s13636-024-00329-7.
C. Shen, L. Zhao, and C. Fu, “Transinger: Cross-Lingual Singing Voice Synthesis via IPA-Based Phonetic Alignment,” Sensors, vol. 25, no. 13, pp. 3973-3991, 2025, doi: 10.3390/s25133973.
Y. Yao, T. Liang, and R. Feng, “SR-TTS: a rhyme-based end-to-end speech synthesis system,” Frontiers in Neurorobotics, vol. 18, no. 1, pp. 1322312-1322322, 2024, doi: 10.3389/fnbot.2024.1322312.
F. Khanam, A. Munmun F, and A. Ritu N, “Text to speech synthesis: a systematic review, deep learning based architecture and future research direction,” Journal of Advances in Information Technology, vol. 13, no. 5, pp. 1-15, 2022, doi: 10.12720/jait.13.5.398-412.
Z. Chen, Z. Ai, and Y. Ma, “Optimizing feature fusion for improved zero-shot adaptation in text-to-speech synthesis,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, pp. 28-45, 2024, doi: 10.1186/s13636-024-00351-9.
C. Galdino J, N. Matos A, and F. Svartman F R, “The evaluation of prosody in speech synthesis: a systematic review,” Journal of the Brazilian Computer Society, vol. 31, no. 1, pp. 465-486, 2025, doi: 10.5753/jbcs.2025.5468.
K. Zhou, B. Sisman, and R. Rana, “Speech synthesis with mixed emotions,” IEEE Transactions on Affective Computing, vol. 14, no. 4, pp. 3120-3134, 2022, doi: 10.1109/TAFFC.2022.3233324.
W. Byun S and P. Lee S, “Design of a multi-condition emotional speech synthesizer,” Applied Sciences, vol. 11, no. 3, pp. 1144-1154, 2021, doi: 10.3390/app11031144.
D. Pawar, P. Borde, and P. Yannawar, “Generating dynamic lip-syncing using target audio in a multimedia environment,” Natural Language Processing Journal, vol. 8, no. 1, pp. 100084-100094, 2024, doi: 10.1016/j.nlp.2024.100084.
Y. Tang, Y. Liu, and W. Li, “Continuous Talking Face Generation Based on Gaussian Blur and Dynamic Convolution,” Sensors, vol. 25, no. 6, pp. 1885-1899, 2025, doi: 10.3390/s25061885.
Z. Wang, W. He, and Y. Wei, “Flow2Flow: Audio-visual cross-modality generation for talking face videos with rhythmic head,” Displays, vol. 80, no. 1, 102552, 2023, doi: 10.1016/j.displa.2023.102552.
Z. Sheng, L. Nie, and M. Liu, “Toward fine-grained talking face generation,” IEEE Transactions on Image Processing, vol. 32, no. 1, pp. 5794-5807, 2023, doi: 10.1109/TIP.2023.3323452.
Z. Liu, X. Liu, and S. Chen, “Multimodal fusion for talking face generation utilizing speech-related facial action units,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 20, no. 9, pp. 1-24, 2024.
Z. Sheng, L. Nie, and M. Zhang, “Stochastic latent talking face generation toward emotional expressions and head poses,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 4, pp. 2734-2748, 2023, doi: 10.1109/TCSVT.2023.3311039.