Application of Conformer Architecture in Clinical Speech Input and Intelligent Medical Record Generation

Main Article Content

X. K. Zou
L. Wang
J. Sun
S. Y. Guo
N. Li

Abstract

Accurate clinical speech recognition remains challenging because rapid pronunciation, domain-specific terminology, and background noise often degrade automatic speech recognition and subsequent medical record generation. This study proposes a multi-stage intelligent documentation framework that integrates a 12-layer Conformer architecture, BERT-BiLSTM-CRF semantic modeling, and BART-based structured text generation. The Conformer encoder captures both local acoustic characteristics and long-range contextual dependencies, while the semantic module performs medical entity recognition and normalization to enhance terminology consistency. The extracted information is subsequently incorporated into a BART generator with clinical knowledge prompts to produce standardized SOAP-compliant medical records. Experimental results demonstrate a word error rate of 6.3%, medical term accuracy of 95.8%, low response latency of approximately 940–960 ms, and generation quality approaching physician-written records. Beyond clinical documentation, the proposed framework illustrates the effectiveness of deep time-frequency feature extraction and contextual sequence modeling for complex noisy signals, offering methodological insights for electromagnetic signal interpretation, antenna measurement data processing, and intelligent information extraction in propagation-related applications.

Downloads

Download data is not yet available.

Article Details

How to Cite
Zou, X. K., Wang, L., Sun, J., Guo, S. Y., & Li, N. (2026). Application of Conformer Architecture in Clinical Speech Input and Intelligent Medical Record Generation. Advanced Electromagnetics, 15(3), 758–768. https://doi.org/10.7716/aem.v15i3.3126
Section
Research Articles

References

Z. T. Guo, W. L. Shi, and T. Yang, “Design and implementation of intelligent voice assistant for clinical diagnosis and treatment of traditional Chinese medicine based on speech recognition,” Journal of Medical Informatics, vol. 43, no. 9, pp. 59-62, 2022, doi: 10.3969/j.issn.1673-6036.2022.09.01.

View Article

J. R. Liu, Y. Song, and R. Jia, “Identification of protected health information in Chinese clinical texts based on BiLSTM-CRF,” Data Analysis and Knowledge Discovery, vol. 4, no. 10, pp. 124-133, 2020, doi: 10.11925/infotech.2096-3467.2020.4.10.124.

View Article

A. S. Miner, A. Haque, J. A. Fries, S. L. Fleming, D. E. Wilfley, G. Terence Wilson, et al., “Assessing the accuracy of automatic speech recognition for psychotherapy,” NPJ Digital Medicine, vol. 3, no. 1, pp. 82, 2020, doi: 10.1038/s41746-020-0285-8.

View Article

Y. Kumar, “A comprehensive analysis of speech recognition systems in healthcare: Current research challenges and future prospects,” SN Computer Science, vol. 5, no. 1, pp. 137, 2024, doi: 10.1007/s42979-023-02466-w.

View Article

L. Sun, A. N. Wang, Y. M. Song, J. Dong, X. L. Liu, H. Liang, et al., “Application, challenges and prospects of large language models in clinical medicine,” Journal of PLA Medical College, vol. 46, no. 1, pp. 50-60, 2025, doi: 10.12435/j.issn.2095-5227.24070201.

View Article

M. Jelassi, O. Jemai, and J. Demongeot, “Revolutionizing radiological analysis: The future of French language automatic speech recognition in healthcare,” Diagnostics, vol. 14, no. 9, pp. 895, 2024, doi: 10.3390/diagnostics14090895.

View Article

B. D. Tran, K. Latif, T. L. Reynolds, and J. Park, ““Mm-hm,”“Uh-uh”: Are non-lexical conversational sounds deal breakers for the ambient clinical documentation technology? Journal of the American Medical Informatics Association,” 2023; 30(4):703-711, doi: 10.1093/jamia/ocad001.

View Article

A. Dutta, G. Ashishkumar, and C. V. R. Rao, “Performance analysis of ASR system in hybrid DNN-HMM framework using a PWL euclidean activation function,” Frontiers of Computer Science, vol. 15, no. 4, Art. no. 154705, 2021, doi: 10.1007/s11704-020-9419-z.

View Article

C. Garcia, D. Nickolai, and L. Jones, “Traditional versus ASR-based pronunciation instruction,” Calico Journal, vol. 37, no. 3, pp. 213-232, 2020, doi: 10.1558/cj.40379.

View Article

A. M. Deshmukh, “Comparison of hidden markov model and recurrent neural network in automatic speech recognition,” European Journal of Engineering and Technology Research, vol. 5, no. 8, pp. 958-965, 2020, doi: 10.24018/eng.2020.5.8.2077.

View Article

F. Ye and J. Yang, “A deep neural network model for speaker identification,” Applied Sciences, vol. 11, no. 8, pp. 3603, 2021, doi: 10.3390/app11083603.

View Article

Z. Ren, N. Yolwas, W. Slamu, R. Cao, and H. Wang, “Improving hybrid ctc/attention architecture for agglutinative language speech recognition,” Sensors, vol. 22, no. 19, pp. 7319, 2022, doi: 10.3390/s22197319.

View Article

L. Tang, “A transformer-based network for speech recognition,” International Journal of Speech Technology, vol. 26, no. 2, pp. 531-539, 2023, doi: 10.1007/s10772-023-10034-z.

View Article

T. Guo, N. Yolwas, and W. Slamu, “Efficient conformer for agglutinative language ASR model using low-rank approximation and balanced softmax,” Applied Sciences, vol. 13, no. 7, pp. 4642, 2023, doi: 10.3390/app13074642.

View Article

P. Jiang, W. Pan, J. Zhang, T. Wang, and J. Huang, “A Robust Conformer-Based Speech Recognition Model for Mandarin Air Traffic Control,” Computers, Materials & Continua, vol. 77, no. 1, pp. 911, 2023, doi: 10.32604/cmc.2023.041772.

View Article

L. Li, Y. Long, D. Xu, and Y. Li, “Boosting Character-based Mandarin ASR via Chinese Pinyin Representation,” International Journal of Speech Technology, vol. 26, no. 4, pp. 895-902, 2023, doi: 10.1007/s10772-023-10050-z.

View Article

Y. Chen, W. H. Ge, and J. Liao, “Research progress of deep learning in named entity recognition in the medical field,” PROGRESS IN PHARMACEUTICAL SCIENCES, vol. 44, no. 1, pp. 28-34, 2020.

W. Zhang, C. Liu, H. B. Fei, W. Li, J. H. Yu, Y. Cao, et al., “Research on speech recognition based on DL-T and transfer learning,” Journal of Engineering Science, vol. 43, no. 3, pp. 433-441, 2021.

X. J. Che, H. Xu, M. Y. Pan, and Q. L. Liu, “Two-stage learning algorithm for biomedical named entity recognition,” Journal of Jilin University (Engineering Edition), vol. 53, no. 8, pp. 2380-2387, 2023, doi: 10.13229/j.cnki.jdxbgxb.20211156.

View Article

R. Mahum, I. Ganiyu, L. Hidri, A. M. El-Sherbeeny, and H. Hassan, “A novel Swin transformer based framework for speech recognition for dysarthria,” Scientific Reports, vol. 15, no. 1, Art. no. 20070, 2025, doi: 10.1038/s41598-025-02042-7.

View Article

X. Yang, A. Chen, N. PourNejatian, H. C. Shin, K. E. Smith, C. Parisien, et al., “A large language model for electronic health records,” NPJ Digital Medicine, vol. 5, no. 1, pp. 194, 2022, doi: 10.1038/s41746-022-00742-2.

View Article

J. Z. Yan, Y. X. He, Z. Y. Luo, H. Hu, S. X. Fan, B. Z. Tang, et al., “Potential typical applications and challenges of generative large language models in the medical field,” Journal of Medical Informatics, vol. 44, no. 9, pp. 23-31, 2023, doi: 10.3969/j.issn.1673-6036.2023.09.003.

View Article

Q. T. Chen, J. W. Ni, J. Xu, X. H. Gao, and L. Z. Xia, “Research on traditional Chinese medicine prescription generation driven by generative artificial intelligence GPT-4,” China Pharmacy, vol. 34, no. 23, pp. 2825-2828, 2023, doi: 10.6039/j.issn.1001-0408.2023.23.02.

View Article

Z. Y. Yang, S. F. Fang, X. C. Chen, and Z. N. Li, “Traditional Chinese medicine question generation based on multistrategy mechanism and BERT,” Journal of Medical Informatics, vol. 43, no. 11, pp. 55-62, 2022, doi: 10.3778/j.issn.1673-9418.2401072.

View Article

T. Mairittha, N. Mairittha, and S. Inoue, “Automatic labeled dialogue generation for nursing record systems,” Journal of Personalized Medicine, vol. 10, no. 3, pp. 62, 2020, doi: 10.3390/jpm10030062.

View Article

J. L. K. E. Fendji, D. C. M. Tala, B. O. Yenke, and M. Atemkeng, “Automatic speech recognition using limited vocabulary: A survey,” Applied Artificial Intelligence, vol. 36, no. 1, Art. no. 2095039, 2022, doi: 10.1080/08839514.2022.2095039.

View Article

C. Tejedor-Garcia, V. Cardenoso-Payo, and D. Escudero-Mancebo, “Automatic speech recognition (ASR) systems applied to pronunciation assessment of L2 Spanish for Japanese speakers,” Applied Sciences, vol. 11, no. 15, pp. 6695, 2021, doi: 10.3390/app11156695.

View Article

Q. Qin, S. Zhao, and C. Liu, “A BERT-BiGRU-CRF model for entity recognition of Chinese electronic medical records,” Complexity, vol. 2021, Art. no. 6631837, 2021, doi: 10.1155/2021/6631837.

View Article

X. Tang, Y. Huang, M. Xia, and C. Long, “A multi-task BERT-BiLSTM-AM-CRF strategy for Chinese named entity recognition,” Neural Processing Letters, vol. 55, no. 2, pp. 1209-1229, 2023, doi: 10.1007/s11063-022-10933-3.

View Article

C. X. Li, D. Li, X. Ling, X. L. Li, J. F. Tang, Y. Zhao, et al., “Discussion on the content and recording standards of the teaching history of Chinese medicine clinical pharmacists,” Chinese Pharmacy, vol. 33, no. 21, pp. 2671-2675, 2023, doi: 10.6039/j.issn.1001-0408.2022.21.20.

View Article

J. Y. Zhai, Y. Lu, S. L. Qian, and D. H. Yu, “Course design of clinical diagnosis and treatment thinking for general practice master’s students of Tongji University,” Chinese General Practice, vol. 26, no. 25, pp. 3202, 2023, doi: 10.12114/j.issn.1007-9572.2022.071.

View Article

W. Yim, Y. Fu, A. Ben Abacha, N. Snider, T. Lin, M. Yetisgen, et al., “Aci-bench: A novel ambient clinical intelligence dataset for benchmarking automatic visit note generation,” Scientific Data, vol. 10, no. 1, pp. 586, 2023, doi: 10.1038/s41597-023-02487-3.

View Article

J. Liang, L. P. Zhang, S. Yan, Y. B. Zhao, and Y. W. Zhang, “Research progress on named entity recognition based on large language model,” Journal of Frontiers of Computer Science & Technology, vol. 18, no. 10, pp. 2594, 2024, doi: 10.3778/j.issn.1673-9418.2407038.

View Article

A. Dhouib, A. Othman, O. El Ghoul, M. K. Khribi, and A. Al Sinani, “Arabic automatic speech recognition: A systematic literature review,” Applied Sciences, vol. 12, no. 17, pp. 8898, 2022, doi: 10.3390/app12178898.

View Article

T. T. N. Ngo, H. H. J. Chen, and K. K. W. Lai, “The effectiveness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis,” ReCALL, vol. 36, no. 1, pp. 4-21, 2024, doi: 10.1017/s0958344023000113.

View Article

Most read articles by the same author(s)

1 2 > >> 

Similar Articles

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > >> 

You may also start an advanced similarity search for this article.