Intelligent Interpretation of Traditional Folk Tales through Vocalization: a Generative Model Based on Emotion-Driven and Lip-Syncing

Main Article Content

S. Li

Abstract

This study presents an integrated multi-modal signal generation framework for intelligent vocalization of traditional folk tales, ensuring emotion-driven speech synthesis with precise lip synchronization. The proposed system models continuous narrative emotion trajectories and embeds them into phoneme-level acoustic generation, using differentiable temporal alignment to synchronize audio and visual outputs. Multi-channel signals—including acoustic features and lip motion vectors— are processed through recursive smoothing and alignment matrices to maintain temporal consistency across modalities. Experimental results demonstrate median phoneme-duration deviations of 19.4–22.8 ms and lip-sync errors within 0.031–0.034, indicating stable audio-visual synchronization under complex narrative rhythms. By interpreting the system as a multi-channel signal processing and closed-loop control framework, the method mirrors principles of precision-engineered communication and control systems, providing insights into temporal alignment, cross-modal consistency, and real-time feedback optimization. This approach offers an engineering-oriented perspective for the development of advanced audio-visual generation, multi-sensor signal integration, and temporally coherent generative models.

Downloads

Download data is not yet available.

Article Details

How to Cite
Li, S. (2026). Intelligent Interpretation of Traditional Folk Tales through Vocalization: a Generative Model Based on Emotion-Driven and Lip-Syncing. Advanced Electromagnetics, 15(3), 2319–2327. https://doi.org/10.7716/aem.v15i3.3285
Section
Research Articles

References

K. Liu, J. Wei, and R. Yuan, “Text-guided dynamic mouth motion capturing for person-generic talking face generation,” Knowledge-Based Systems, vol. 316, no. 1, pp. 113354-113380, 2025, doi: 10.1016/j.knosys.2025.113354.

View Article

X. Wang, Y. Huo, and Y. Liu, “Multimodal Feature-Guided Audio-Driven Emotional Talking Face Generation,” Electronics, vol. 14, no. 13, pp. 2684-2704, 2025, doi: 10.3390/electronics14132684.

View Article

J. Wang, O. Zhang, and Y. Jiang, “Multimodal diffusion framework for collaborative text image audio generation and applications,” Scientific Reports, vol. 15, no. 1, pp. 20604-20618, 2025, doi: 10.1038/s41598-025-05794-4.

View Article

J. Xu, X. Wu, and Z. Zhang, “MARS: Multimodal-Assisted Refined Semantic Alignment,” Information Processing & Management, vol. 63, no. 1, 104292, 2026, doi: 10.1016/j.ipm.2025.104292.

View Article

D. Jiang, J. Chang, and L. You, “Audio-Driven Facial Animation with Deep Learning: A Survey,” Information, vol. 15, no. 11, pp. 675-698, 2024, doi: 10.3390/info15110675.

View Article

X. Ji, Z. Liao, and L. Dong, “3D facial animation driven by speech-video dual-modal signals,” Complex & Intelligent Systems, vol. 10, no. 5, pp. 5951-5964, 2024, doi: 10.1007/s40747-024-01481-5.

View Article

L. Yu, H. Xie, and Y. Zhang, “Multimodal learning for temporally coherent talking face generation with articulator synergy,” IEEE Transactions on Multimedia, vol. 24, no. 1, pp. 2950-2962, 2021.

J. Salas-Caceres, J. Lorenzo-Navarro, and D. Freire-Obregon, “Multimodal emotion recognition based on a fusion of audio visual information with temporal dynamics,” Multimedia tools and applications, vol. 84, no. 23, pp. 27327-27343, 2025, doi: 10.1007/s11042-024-20227-6.

View Article

S. Liang, R. Zhou, and Q. Yuan, “ECE-TTS: A Zero-Shot Emotion Text-to-Speech Model with Simplified and Precise Control,” Applied Sciences, vol. 15, no. 9, pp. 5108-5129, 2025, doi: 10.3390/app15095108.

View Article

S. Hwang J, H. Lee S, and W. Lee S, “HiddenSinger: High-quality singing voice synthesis via neural audio co-dec and latent diffusion models,” Neural Networks, vol. 181, no. 1, pp. 106762-106771, 2025, doi: 10.1016/j.neunet.2024.106762.

View Article

R. Wu, Y. Yu, and F. Zhan, “Audio-driven talking face generation with diverse yet realistic facial animations,” Pattern Recognition, vol. 144, no. 1, pp. 109865-109891, 2023, doi: 10.1016/j.patcog.2023.109865.

View Article

X. Luo, S. Takamichi, and Y. Saito, “Emotion-controllable Speech Synthesis Using Emotion Soft Label, Utterance-level Prosodic Factors, and Word-level Prominence,” APSIPA Transactions on Signal and Information Processing, vol. 13, no. 1, pp. 1-30, 2024, doi: 10.1561/116.00000242.

View Article

W. Zhao and Z. Yang, “An emotion speech synthesis method based on vits,” Applied Sciences, vol. 13, no. 4, pp. 2225-2236, 2023, doi: 10.3390/app13042225.

View Article

Y. Hono, K. Hashimoto, and K. Oura, “Sinsy: A deep neural network-based singing voice synthesis system,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, no. 1, pp. 2803-2815, 2021, doi: 10.1109/TASLP.2021.3104165.

View Article

Y. Zhao, M. Yang, and Y. Lin, “AI-enabled text-to-music generation: A comprehensive review of methods, frameworks, and future directions,” Electronics, vol. 14, no. 6, pp. 1197-1249, 2025, doi: 10.3390/electronics14061197.

View Article

L. Liu, J. Wang, and S. Chen, “VividWav2Lip: high-fidelity facial animation generation based on speech-driven lip synchronization,” Electronics, vol. 13, no. 18, pp. 3657-3675, 2024, doi: 10.3390/electronics13183657.

View Article

Y. Fan, Z. Lin, and J. Saito, “Joint audio-text model for expressive speech-driven 3d facial animation,” Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 5, no. 1, pp. 1-15, 2022, doi: 10.1145/3522615.

View Article

X. Liu, X. Liu, and P. Yang, “An approach to optimizing semantic consistency for text-to-digital human generation,” Engineering Applications of Artificial Intelligence, vol. 160, no. 1, 111909, 2025, doi: 10.1016/j.engappai.2025.111909.

View Article

G. Chen, X. Liang, and W. He, “Enhancing Human-Computer Interaction Through Decoupling Motion and Camera Control in Human-Centric Video Generation,” International Journal of Human-Computer Interaction, vol. 1, no. 1, pp. 1-15, 2025, doi: 10.1080/10447318.2025.2487877.

View Article

J. Yang, J. Liu, and K. Huang, “Single-and cross-lingual speech emotion recognition based on WavLM domain emotion embedding,” Electronics, vol. 13, no. 7, pp. 1380-1395, 2024, doi: 10.3390/electronics13071380.

View Article

T. Lodewyckx, F. Tuerlinckx, and P. Kuppens, “A hierarchical state space approach to affective dynamics,” Journal of mathematical psychology, vol. 55, no. 1, pp. 68-83, 2011, doi: 10.1016/j.jmp.2010.08.004.

View Article

S. Sadok, S. Leglaive, and L. Girin, “A multimodal dynamical variational autoencoder for audiovisual speech representation learning,” Neural Networks, vol. 172, no. 1, pp. 106120-106135, 2024, doi: 10.1016/j.neunet.2024.106120.

View Article

H. Barakat, O. Turk, and C. Demiroglu, “Deep learning-based expressive speech synthesis: a systematic review of approaches, challenges, and resources,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, pp. 11-44, 2024, doi: 10.1186/s13636-024-00329-7.

View Article

C. Shen, L. Zhao, and C. Fu, “Transinger: Cross-Lingual Singing Voice Synthesis via IPA-Based Phonetic Alignment,” Sensors, vol. 25, no. 13, pp. 3973-3991, 2025, doi: 10.3390/s25133973.

View Article

Y. Yao, T. Liang, and R. Feng, “SR-TTS: a rhyme-based end-to-end speech synthesis system,” Frontiers in Neurorobotics, vol. 18, no. 1, pp. 1322312-1322322, 2024, doi: 10.3389/fnbot.2024.1322312.

View Article

F. Khanam, A. Munmun F, and A. Ritu N, “Text to speech synthesis: a systematic review, deep learning based architecture and future research direction,” Journal of Advances in Information Technology, vol. 13, no. 5, pp. 1-15, 2022, doi: 10.12720/jait.13.5.398-412.

View Article

Z. Chen, Z. Ai, and Y. Ma, “Optimizing feature fusion for improved zero-shot adaptation in text-to-speech synthesis,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, pp. 28-45, 2024, doi: 10.1186/s13636-024-00351-9.

View Article

C. Galdino J, N. Matos A, and F. Svartman F R, “The evaluation of prosody in speech synthesis: a systematic review,” Journal of the Brazilian Computer Society, vol. 31, no. 1, pp. 465-486, 2025, doi: 10.5753/jbcs.2025.5468.

View Article

K. Zhou, B. Sisman, and R. Rana, “Speech synthesis with mixed emotions,” IEEE Transactions on Affective Computing, vol. 14, no. 4, pp. 3120-3134, 2022, doi: 10.1109/TAFFC.2022.3233324.

View Article

W. Byun S and P. Lee S, “Design of a multi-condition emotional speech synthesizer,” Applied Sciences, vol. 11, no. 3, pp. 1144-1154, 2021, doi: 10.3390/app11031144.

View Article

D. Pawar, P. Borde, and P. Yannawar, “Generating dynamic lip-syncing using target audio in a multimedia environment,” Natural Language Processing Journal, vol. 8, no. 1, pp. 100084-100094, 2024, doi: 10.1016/j.nlp.2024.100084.

View Article

Y. Tang, Y. Liu, and W. Li, “Continuous Talking Face Generation Based on Gaussian Blur and Dynamic Convolution,” Sensors, vol. 25, no. 6, pp. 1885-1899, 2025, doi: 10.3390/s25061885.

View Article

Z. Wang, W. He, and Y. Wei, “Flow2Flow: Audio-visual cross-modality generation for talking face videos with rhythmic head,” Displays, vol. 80, no. 1, 102552, 2023, doi: 10.1016/j.displa.2023.102552.

View Article

Z. Sheng, L. Nie, and M. Liu, “Toward fine-grained talking face generation,” IEEE Transactions on Image Processing, vol. 32, no. 1, pp. 5794-5807, 2023, doi: 10.1109/TIP.2023.3323452.

View Article

Z. Liu, X. Liu, and S. Chen, “Multimodal fusion for talking face generation utilizing speech-related facial action units,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 20, no. 9, pp. 1-24, 2024.

Z. Sheng, L. Nie, and M. Zhang, “Stochastic latent talking face generation toward emotional expressions and head poses,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 4, pp. 2734-2748, 2023, doi: 10.1109/TCSVT.2023.3311039.

View Article

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > >> 

You may also start an advanced similarity search for this article.