Precise Guidance for English Prosody Perception Pronunciation Based on Deep Learning
Main Article Content
Abstract
In English pronunciation teaching, prosody has long received less attention than segmental phoneme training, although stress, rhythm, and intonation strongly affect speech intelligibility and naturalness. Existing technological aids mainly focus on phoneme-level correction and still lack fine-grained diagnosis and adaptive feedback for prosodic waveform patterns. From an engineering perspective, prosody perception can be regarded as a speech-signal analysis problem involving pitch trajectory, energy distribution, temporal modulation, and waveform feature recognition, which is methodologically related to time-frequency analysis in wave-based signal processing. This study constructs an intelligent perception and precise guidance system for English prosody based on deep learning. Speech samples were collected from English learners with different dialect backgrounds, and a specialized corpus containing 4592 valid speech samples was constructed. Acoustic features including fundamental frequency, duration, intensity, energy rising points, and intonation trajectories were extracted. A classifier ensemble model was trained for stress detection, machine learning models were used for rhythm feature analysis, and a convolutional neural network was developed for intonation pattern recognition. The system further integrates diagnostic results into visual feedback, comparative listening, and graded practice modules. The results show that the proposed models can effectively identify learners’ prosodic errors across stress, rhythm, and intonation dimensions. The teaching experiment indicates that the system significantly improves learners’ prosody perception and pronunciation naturalness, providing a data-driven technical route for intelligent and precise oral pronunciation guidance.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
E. Akhter, “The Impact of Human-Machine Interaction On English Pronunciation And Fluency: Case Studies Using AI Speech Assistants,” Review of Applied Science and Technology, vol. 4, no. 02, pp. 473-500, 2025, doi: 10.63125/1wyj3p84.
N. T. Hoang, D. N. Han, and D. H. Le, “Exploring Chatbot AI in improving vocational students English pronunciation,” AsiaCALL Online Journal, vol. 14, no. 2, pp. 140-155, 2023, doi: 10.54855/acoj.231429.
M. Bashori, R. van Hout, H. Strik, et al., “I Can Speak: improving English pronunciation through automatic speech recognition-based language learning systems,” Innovation in Language Learning and Teaching, vol. 18, no. 5, pp. 443-461, 2024, doi: 10.1080/17501229.2024.2315101.
R. P. Octaviani, L. M. Jannah, M. Sebrina, et al., “The impacts of first language on students English pronunciation,” IJIET (International Journal of Indonesian Education and Teaching), vol. 8, no. 1, pp. 164-173, 2024, doi: 10.24071/ijiet.v8i1.6758.
H. S. Utami and R. Morganna, “Improving students English pronunciation competence by using shadowing technique,” English Franca: Academic Journal of English Language and Education, vol. 6, no. 1, pp. 127-150, 2022, doi: 10.29240/ef.v6i1.3915.
V. G. Sardegna, “Evidence in favor of a strategy-based model for English pronunciation instruction,” Language Teaching, vol. 55, no. 3, pp. 363-378, 2022, doi: 10.1017/S0261444821000380.
I. Iskandar, R. Dewanti, S. D. Sulistyaningrum, et al., “Scaffolding Assignments to Conciliate the Disinclination to Employ Project-Based Learning of English Pronunciation and Autodidacticism,” International Journal of Language Education, vol. 8, no. 2, pp. 199-227, 2024, doi: 10.26858/ijole.v8i2.64087.
W. Dandee and P. Pornwiriyakit, “Improving English Pronunciation Skills by Using English Phonetic Alphabet Drills in EFL Students,” Journal of Educational Issues, vol. 8, no. 1, pp. 611-628, 2022, doi: 10.5296/jei.v8i1.19851.
W. M. Hsieh, H. C. Yeh, and N. S. Chen, “Impact of a robot and tangible object (R&T) integrated learning system on elementary EFL learners English pronunciation and willingness to communicate,” Computer Assisted Language Learning, vol. 38, no. 4, pp. 773-798, 2025, doi: 10.1080/09588221.2023.2228357.
A. Almusharraf, “Pronunciation instruction in the context of world English: exploring university EFL instructors perceptions and practices,” Humanities and Social Sciences Communications, vol. 11, no. 1, pp. 1-11, 2024, doi: 10.1057/s41599-024-03365-y.
F. M. A. M. Awadh, N. H. P. S. Putro, A. E. Pohan, et al., “Improving English pronunciation through phonetics instruction in Yemeni EFL classrooms,” JOLLT Journal of Languages and Language Teaching, vol. 12, no. 2, pp. 930-940, 2024, doi: 10.33394/jollt.v12i2.10720.
A. D. J. Rubio, “The situation of English pronunciation in primary education classrooms in Spain,” International Journal of Instruction, vol. 17, no. 3, pp. 529-544, 2024, doi: 10.29333/iji.2024.17329a.
M. D. Ammar, “English pronunciation problems analysis faced by English education students in the second semester at Indo Global Mandiri University,” Global Expert: Jurnal Bahasa Dan Sastra, vol. 10, no. 1, pp. 1-7, 2022, doi: 10.36982/jge.v10i1.2166.
X. Xue and R. E. Dunham, “Using a SPOC-based flipped classroom instructional mode to teach English pronunciation,” Computer Assisted Language Learning, vol. 36, no. 7, pp. 1309-1337, 2023, doi: 10.1080/09588221.2021.1980404.