Development of Real-Time Oral Error Correction System for College English Classrooms Based on BERT
Main Article Content
Abstract
This paper presents a real-time oral error correction system for college English classrooms based on an acoustic-semantic fusion DistilBERT+Adapter architecture. Whisper-small is used for speech transcription, and ASR confidence scores and word-duration features are embedded directly into the BERT representation space to improve robustness against speech-recognition noise. The model jointly performs error localization through a CRF layer and error-type classification, and the resulting outputs guide a constrained decoding mechanism that generates Top-3 correction candidates. These candidates are subsequently re-ranked using a KenLM language model. The system is lightweight and efficient, containing only 44M parameters and achieving an inference latency of 190 ms. End-to-end evaluation shows a latency of 438 ± 52 ms, Accuracy@Top1 of 73.1%, F0.5 of 0.692, and a teacher rating of 4.2. Through adapter fine-tuning, knowledge distillation, and ONNX runtime optimization, the proposed system achieves strong noise robustness and generalization, offering a deployable solution for personalized oral English instruction and real-time acoustic-semantic signal processing.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
F. You, “Exploration of the “scaffolding” teaching model based on information technology in oral English classes,” Development Pedagogy (Chinese Version), vol. 3, no. 7, pp. 64-66, 2022, doi: 10.12184/wspfzjyxWSP2634-793818.20220307.
J. Xu and C. Zhao, “The role of large language models in English teaching,” Frontiers of Foreign Language Education in China, vol. 7, no. 1, pp. 3-10, 2024, doi: 10.20083/j.cnki.fleic.2024.01.003.
E. Kartchava, E. Gatbonton, A. Ammar, and P. Trofimovich, “Oral corrective feedback: Pre-service English as a second language teachers’ beliefs and practices,” Language Teaching Research, vol. 24, no. 2, pp. 220-249, 2020, doi: 10.1177/1362168818787546.
A. Kholis, “Elsa speak app: Automatic speech recognition (ASR) for supplementing English pronunciation skills,” Pedagogy: Journal of English Language Teaching, vol. 9, no. 1, pp. 1-14, 2021, doi: 10.32332/joelt.v13i2.
S. Zhou and W. Liu, “English grammar error correction algorithm based on classification model,” Complexity, vol. 2021, no. 1, Art. no. 6687337, 2021, doi: 10.1155/2021/6687337.
S. McCrocklin and I. Edalatishams, “Revisiting popular speech recognition software for ESL speech,” Tesol Quarterly, vol. 54, no. 4, pp. 1086-1097, 2020, [Online]. Available: https://www.jstor.org/stable/48642909.
M. Y. C. Jiang, M. S. Y. Jong, W. W. F. Lau, C. S. Chai, and N. Wu, “Exploring the effects of automatic speech recognition technology on oral accuracy and fluency in a flipped classroom,” Journal of Computer Assisted Learning, vol. 39, no. 1, pp. 125-140, 2023, doi: 10.1111/jcal.12732.
L. Cheng, P. Ben, and Y. Qiao, “Research on automatic error correction method in English writing based on deep neural network,” Computational Intelligence and Neuroscience, vol. 2022, no. 1, Art. no. 2709255, 2022, doi: 10.1155/2022/2709255.
H. U. Sakiroglu, “Oral Corrective Feedback Preferences of University Students in English Communication Classes,” International Journal of Research in Education and Science, vol. 6, no. 1, pp. 172-178, 2020, doi: 10.46328/ijres.v6i1.806.
T. T. N. Ngo, H. H. J. Chen, and K. K. W. Lai, “The effectiveness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis,” ReCALL, vol. 36, no. 1, pp. 4-21, 2024, doi: 10.1017/S0958344023000113.
K. Evers and S. Chen, “Effects of an automatic speech recognition system with peer feedback on pronunciation instruction for adults,” Computer Assisted Language Learning, vol. 35, no. 8, pp. 1869-1889, 2022, doi: 10.1080/09588221.2020.1839504.
Y. Qing, “Design and application of automatic English translation grammar error detection system based on bert machine vision,” Scalable Computing: Practice and Experience, vol. 25, no. 3, pp. 2088-2102, 2024, doi: 10.12694/scpe.v25i3.2770.
R. Raju, P. B. Pati, S. A. Gandheesh, G. S. Sannala, and K. S. Suriya, “Grammatical versus spelling error correction: An investigation into the responsiveness of transformer-based language models using BART and MarianMT,” Journal of Information & Knowledge Management, vol. 23, no. 3, Art. no. 2450037, 2024, doi: 10.1142/S0219649224500370.
M. S. I. Malik, A. Nazarova, M. M. Jamjoom, and D. I. Ignatov, “Multilingual hope speech detection: A Robust framework using transfer learning of fine-tuning RoBERTa model,” Journal of King Saud University - Computer and Information Sciences, vol. 35, no. 8, Art. no. 101736, 2023, doi: 10.1016/j.jksuci.2023.101736.
M. C. Tudose, S. Ruseti, and M. Dascalu, “Show Me All Writing Errors: A Two-Phased Grammatical Error Corrector for Romanian,” Information, vol. 16, no. 3, pp. 242, 2025, doi: 10.3390/info16030242.