Application of Conformer Architecture in Clinical Speech Input and Intelligent Medical Record Generation
Main Article Content
Abstract
Accurate clinical speech recognition remains challenging because rapid pronunciation, domain-specific terminology, and background noise often degrade automatic speech recognition and subsequent medical record generation. This study proposes a multi-stage intelligent documentation framework that integrates a 12-layer Conformer architecture, BERT-BiLSTM-CRF semantic modeling, and BART-based structured text generation. The Conformer encoder captures both local acoustic characteristics and long-range contextual dependencies, while the semantic module performs medical entity recognition and normalization to enhance terminology consistency. The extracted information is subsequently incorporated into a BART generator with clinical knowledge prompts to produce standardized SOAP-compliant medical records. Experimental results demonstrate a word error rate of 6.3%, medical term accuracy of 95.8%, low response latency of approximately 940–960 ms, and generation quality approaching physician-written records. Beyond clinical documentation, the proposed framework illustrates the effectiveness of deep time-frequency feature extraction and contextual sequence modeling for complex noisy signals, offering methodological insights for electromagnetic signal interpretation, antenna measurement data processing, and intelligent information extraction in propagation-related applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
Z. T. Guo, W. L. Shi, and T. Yang, “Design and implementation of intelligent voice assistant for clinical diagnosis and treatment of traditional Chinese medicine based on speech recognition,” Journal of Medical Informatics, vol. 43, no. 9, pp. 59-62, 2022, doi: 10.3969/j.issn.1673-6036.2022.09.01.
J. R. Liu, Y. Song, and R. Jia, “Identification of protected health information in Chinese clinical texts based on BiLSTM-CRF,” Data Analysis and Knowledge Discovery, vol. 4, no. 10, pp. 124-133, 2020, doi: 10.11925/infotech.2096-3467.2020.4.10.124.
A. S. Miner, A. Haque, J. A. Fries, S. L. Fleming, D. E. Wilfley, G. Terence Wilson, et al., “Assessing the accuracy of automatic speech recognition for psychotherapy,” NPJ Digital Medicine, vol. 3, no. 1, pp. 82, 2020, doi: 10.1038/s41746-020-0285-8.
Y. Kumar, “A comprehensive analysis of speech recognition systems in healthcare: Current research challenges and future prospects,” SN Computer Science, vol. 5, no. 1, pp. 137, 2024, doi: 10.1007/s42979-023-02466-w.
L. Sun, A. N. Wang, Y. M. Song, J. Dong, X. L. Liu, H. Liang, et al., “Application, challenges and prospects of large language models in clinical medicine,” Journal of PLA Medical College, vol. 46, no. 1, pp. 50-60, 2025, doi: 10.12435/j.issn.2095-5227.24070201.
M. Jelassi, O. Jemai, and J. Demongeot, “Revolutionizing radiological analysis: The future of French language automatic speech recognition in healthcare,” Diagnostics, vol. 14, no. 9, pp. 895, 2024, doi: 10.3390/diagnostics14090895.
B. D. Tran, K. Latif, T. L. Reynolds, and J. Park, ““Mm-hm,”“Uh-uh”: Are non-lexical conversational sounds deal breakers for the ambient clinical documentation technology? Journal of the American Medical Informatics Association,” 2023; 30(4):703-711, doi: 10.1093/jamia/ocad001.
A. Dutta, G. Ashishkumar, and C. V. R. Rao, “Performance analysis of ASR system in hybrid DNN-HMM framework using a PWL euclidean activation function,” Frontiers of Computer Science, vol. 15, no. 4, Art. no. 154705, 2021, doi: 10.1007/s11704-020-9419-z.
C. Garcia, D. Nickolai, and L. Jones, “Traditional versus ASR-based pronunciation instruction,” Calico Journal, vol. 37, no. 3, pp. 213-232, 2020, doi: 10.1558/cj.40379.
A. M. Deshmukh, “Comparison of hidden markov model and recurrent neural network in automatic speech recognition,” European Journal of Engineering and Technology Research, vol. 5, no. 8, pp. 958-965, 2020, doi: 10.24018/eng.2020.5.8.2077.
F. Ye and J. Yang, “A deep neural network model for speaker identification,” Applied Sciences, vol. 11, no. 8, pp. 3603, 2021, doi: 10.3390/app11083603.
Z. Ren, N. Yolwas, W. Slamu, R. Cao, and H. Wang, “Improving hybrid ctc/attention architecture for agglutinative language speech recognition,” Sensors, vol. 22, no. 19, pp. 7319, 2022, doi: 10.3390/s22197319.
L. Tang, “A transformer-based network for speech recognition,” International Journal of Speech Technology, vol. 26, no. 2, pp. 531-539, 2023, doi: 10.1007/s10772-023-10034-z.
T. Guo, N. Yolwas, and W. Slamu, “Efficient conformer for agglutinative language ASR model using low-rank approximation and balanced softmax,” Applied Sciences, vol. 13, no. 7, pp. 4642, 2023, doi: 10.3390/app13074642.
P. Jiang, W. Pan, J. Zhang, T. Wang, and J. Huang, “A Robust Conformer-Based Speech Recognition Model for Mandarin Air Traffic Control,” Computers, Materials & Continua, vol. 77, no. 1, pp. 911, 2023, doi: 10.32604/cmc.2023.041772.
L. Li, Y. Long, D. Xu, and Y. Li, “Boosting Character-based Mandarin ASR via Chinese Pinyin Representation,” International Journal of Speech Technology, vol. 26, no. 4, pp. 895-902, 2023, doi: 10.1007/s10772-023-10050-z.
Y. Chen, W. H. Ge, and J. Liao, “Research progress of deep learning in named entity recognition in the medical field,” PROGRESS IN PHARMACEUTICAL SCIENCES, vol. 44, no. 1, pp. 28-34, 2020.
W. Zhang, C. Liu, H. B. Fei, W. Li, J. H. Yu, Y. Cao, et al., “Research on speech recognition based on DL-T and transfer learning,” Journal of Engineering Science, vol. 43, no. 3, pp. 433-441, 2021.
X. J. Che, H. Xu, M. Y. Pan, and Q. L. Liu, “Two-stage learning algorithm for biomedical named entity recognition,” Journal of Jilin University (Engineering Edition), vol. 53, no. 8, pp. 2380-2387, 2023, doi: 10.13229/j.cnki.jdxbgxb.20211156.
R. Mahum, I. Ganiyu, L. Hidri, A. M. El-Sherbeeny, and H. Hassan, “A novel Swin transformer based framework for speech recognition for dysarthria,” Scientific Reports, vol. 15, no. 1, Art. no. 20070, 2025, doi: 10.1038/s41598-025-02042-7.
X. Yang, A. Chen, N. PourNejatian, H. C. Shin, K. E. Smith, C. Parisien, et al., “A large language model for electronic health records,” NPJ Digital Medicine, vol. 5, no. 1, pp. 194, 2022, doi: 10.1038/s41746-022-00742-2.
J. Z. Yan, Y. X. He, Z. Y. Luo, H. Hu, S. X. Fan, B. Z. Tang, et al., “Potential typical applications and challenges of generative large language models in the medical field,” Journal of Medical Informatics, vol. 44, no. 9, pp. 23-31, 2023, doi: 10.3969/j.issn.1673-6036.2023.09.003.
Q. T. Chen, J. W. Ni, J. Xu, X. H. Gao, and L. Z. Xia, “Research on traditional Chinese medicine prescription generation driven by generative artificial intelligence GPT-4,” China Pharmacy, vol. 34, no. 23, pp. 2825-2828, 2023, doi: 10.6039/j.issn.1001-0408.2023.23.02.
Z. Y. Yang, S. F. Fang, X. C. Chen, and Z. N. Li, “Traditional Chinese medicine question generation based on multistrategy mechanism and BERT,” Journal of Medical Informatics, vol. 43, no. 11, pp. 55-62, 2022, doi: 10.3778/j.issn.1673-9418.2401072.
T. Mairittha, N. Mairittha, and S. Inoue, “Automatic labeled dialogue generation for nursing record systems,” Journal of Personalized Medicine, vol. 10, no. 3, pp. 62, 2020, doi: 10.3390/jpm10030062.
J. L. K. E. Fendji, D. C. M. Tala, B. O. Yenke, and M. Atemkeng, “Automatic speech recognition using limited vocabulary: A survey,” Applied Artificial Intelligence, vol. 36, no. 1, Art. no. 2095039, 2022, doi: 10.1080/08839514.2022.2095039.
C. Tejedor-Garcia, V. Cardenoso-Payo, and D. Escudero-Mancebo, “Automatic speech recognition (ASR) systems applied to pronunciation assessment of L2 Spanish for Japanese speakers,” Applied Sciences, vol. 11, no. 15, pp. 6695, 2021, doi: 10.3390/app11156695.
Q. Qin, S. Zhao, and C. Liu, “A BERT-BiGRU-CRF model for entity recognition of Chinese electronic medical records,” Complexity, vol. 2021, Art. no. 6631837, 2021, doi: 10.1155/2021/6631837.
X. Tang, Y. Huang, M. Xia, and C. Long, “A multi-task BERT-BiLSTM-AM-CRF strategy for Chinese named entity recognition,” Neural Processing Letters, vol. 55, no. 2, pp. 1209-1229, 2023, doi: 10.1007/s11063-022-10933-3.
C. X. Li, D. Li, X. Ling, X. L. Li, J. F. Tang, Y. Zhao, et al., “Discussion on the content and recording standards of the teaching history of Chinese medicine clinical pharmacists,” Chinese Pharmacy, vol. 33, no. 21, pp. 2671-2675, 2023, doi: 10.6039/j.issn.1001-0408.2022.21.20.
J. Y. Zhai, Y. Lu, S. L. Qian, and D. H. Yu, “Course design of clinical diagnosis and treatment thinking for general practice master’s students of Tongji University,” Chinese General Practice, vol. 26, no. 25, pp. 3202, 2023, doi: 10.12114/j.issn.1007-9572.2022.071.
W. Yim, Y. Fu, A. Ben Abacha, N. Snider, T. Lin, M. Yetisgen, et al., “Aci-bench: A novel ambient clinical intelligence dataset for benchmarking automatic visit note generation,” Scientific Data, vol. 10, no. 1, pp. 586, 2023, doi: 10.1038/s41597-023-02487-3.
J. Liang, L. P. Zhang, S. Yan, Y. B. Zhao, and Y. W. Zhang, “Research progress on named entity recognition based on large language model,” Journal of Frontiers of Computer Science & Technology, vol. 18, no. 10, pp. 2594, 2024, doi: 10.3778/j.issn.1673-9418.2407038.
A. Dhouib, A. Othman, O. El Ghoul, M. K. Khribi, and A. Al Sinani, “Arabic automatic speech recognition: A systematic literature review,” Applied Sciences, vol. 12, no. 17, pp. 8898, 2022, doi: 10.3390/app12178898.
T. T. N. Ngo, H. H. J. Chen, and K. K. W. Lai, “The effectiveness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis,” ReCALL, vol. 36, no. 1, pp. 4-21, 2024, doi: 10.1017/s0958344023000113.