Fusion of Improved BERT and BiLSTM Models to Improve English-Chinese Terminology Translation
Main Article Content
Abstract
To address weak domain adaptability and insufficient modeling of long-range contextual dependencies in specialized English-Chinese terminology translation, this paper proposes a neural translation model integrating improved BERT and BiLSTM. The model first pre-trains BERT on the public WMT corpus and domain-specific corpora, and introduces term boundary markers to enhance input representation, improving recognition of specialized terminology and domain adaptability. A gated-attention BiLSTM encoder is then designed to strengthen local contextual awareness and sequence-structure modeling, especially around term boundaries, thereby complementing BERT’s global representation and capturing the semantic context of terms more accurately. A feature-fusion module integrates multi-layer representations from BERT and BiLSTM, and an attention-driven decoder generates the final translation. Experimental results show that the proposed model outperforms advanced baselines across key metrics. Terminology translation accuracy reaches 88.1%, domain consistency reaches 92.0%, term-boundary F1-score reaches 87.6%, and context sensitivity reaches 86.5%, verifying the effectiveness of the method in improving the accuracy and robustness of specialized terminology translation.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
H. Wang, H. Wu, Z. He, L. Huang, and K. W. Church, “Progress in machine translation,” Engineering, vol. 18, pp. 143-153, 2022, doi: 10.1016/j.eng.2021.03.023.
P. Naveen and P. Trojovský, “Overview and challenges of machine translation for contextually appropriate translations,” Iscience, vol. 27, no. 10, 110856, 2024, doi: 10.1016/j.isci.2024.110878.
R. Al-Jarf, “Translation of Arabic Folk Medical Terms with Om and Abu by AI: A Comparison of Microsoft Copilot and DeepSeek,” Journal of Medical and Health Studies, vol. 6, no. 4, pp. 45-58, 2025, doi: 10.32996/jmhs.2025.6.4.8.
X. E. Zhang, “A study of cultural context in Chinese-English translation,” Region-Educational Research and Reviews, vol. 3, no. 2, pp. 11-14, 2021, doi: 10.32629/rerr.v3i2.303.
S. A. Mohamed, A. A. Elsayed, Y. F. Hassan, and M. A. Abdou, “Neural machine translation: past, present, and future,” Neural Computing and Applications, vol. 33, no. 23, pp. 15919-15931. doi: 10.1007/s00521-021-06268-0, 2021.
G. Tang, “Translation of the proper nouns in legal English,” Theory and Practice in Language Studies, vol. 11, no. 6, pp. 711-716, 2021, doi: 10.17507/tpls.1106.15.
Guerberof ArenasA, J. Moorkens, and S. O’Brien, “The impact of translation modality on user experience: an eye-tracking study of the Microsoft Word user interface,” Machine Translation, vol. 35, no. 2, pp. 205-237, 2021, doi: 10.1007/s10590-021-09267-z.
H. Wang, J. Li, H. Wu, E. Hovy, and Y. Sun, “Pre-trained language models and their applications,” Engineering, vol. 25, pp. 51-65, 2023, doi: 10.1016/j.eng.2022.04.024.
B. Min, H. Ross, E. Sulem, et al., “Recent advances in natural language processing via large pre-trained language models: A survey,” ACM Computing Surveys, vol. 56, no. 2, pp. 1-40, 2023, doi: 10.1145/3605943.
É. Arnaud, M. Elbattah, M. Gignon, and D. Gilles, “Learning Embeddings from Free-text Triage Notes using Pretrained Transformer Models,” HEALTHINF, vol. 5, pp. 835-841, 2022, doi: 10.5220/0011012800003123.
Y. Zhu, “A knowledge graph and BiLSTM-CRF-enabled intelligent adaptive learning model and its potential application,” Alexandria Engineering Journal, vol. 91, pp. 305-320, 2024, doi: 10.1016/j.aej.2024.02.011.
S. Arslan, “Application of BiLSTM-CRF model with different embeddings for product name extraction in unstructured Turkish text,” Neural Computing and Applications, vol. 36, no. 15, pp. 8371-8382, 2024, doi: 10.1007/s00521-024-09532-1.
M. M. Rahman, A. I. Shiplu, Y. Watanobe, and M. A. Alam, “RoBERTa-BiLSTM: A context-aware hybrid model for sentiment analysis,” IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 9, pp. 1-12, 2025, doi: 10.1109/TETCI.2025.3572150.
Y. Cui and M. Liang, “Automated scoring of translations with BERT models: Chinese and English language case study,” Applied Sciences, vol. 14, no. 5, pp. 1925-1925, 2024, doi: 10.3390/app14051925.
S. Doddapaneni, G. Ramesh, M. Khapra, A. Kunchukuttan, and P. Kumar, “A primer on pretrained multilingual language models,” ACM Computing Surveys, vol. 57, no. 9, pp. 1-39, 2025, doi: 10.1145/3727339.
M. Jiralerspong, J. Bose, I. Gemp, C. Qin, Y. Bachrach, and G. Gidel, “Feature likelihood divergence: evaluating the generalization of generative models using samples,” Advances in Neural Information Processing Systems, vol. 36, pp. 33095-33119, 2023, doi: 10.52202/075280-1436.
P. Zhou, X. Xie, Z. Lin, and S. Yan, “Towards understanding convergence and generalization of AdamW,” IEEE transactions on pattern analysis and machine intelligence, vol. 46, no. 9, pp. 6486-6493, 2024, doi: 10.1109/TPAMI.2024.3382294.
Z. Ren and Z. Zhou, “Dynamic batch learning in high-dimensional sparse linear contextual bandits,” Management Science, vol. 70, no. 2, pp. 1315-1342, 2024, doi: 10.1287/mnsc.2023.4895.
A. Jolly, V. Pandey, I. Singh, and N. Sharma, “Exploring biomedical named entity recognition via SciSpaCy and BioBERT models,” The Open Biomedical Engineering Journal, vol. 18, no. 1, pp. 1-10, 2024, doi: 10.2174/0118741207289680240510045617.
B. Ranjgar, A. Sadeghi-Niaraki, M. Shakeri, F. Rahimi, and S. Choi, “Cultural heritage information retrieval: past, present, and future trends,” IEEE Access, vol. 12, pp. 42992-43026, 2024, doi: 10.1109/ACCESS.2024.3374769.
C. Si, Z. Zhang, Y. Chen, F. Qi, X. Wang, Z. Liu, Y. Wang, and et.al, “Sub-character tokenization for Chinese pretrained language models,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 469-487, 2023, doi: 10.1162/tacl_a_00560.
A. Hernández and J. M. Amigó, “Attention mechanisms and their applications to complex systems,” Entropy, vol. 23, no. 3, pp. 118-147, 2021, doi: 10.3390/e23030283.
S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Mariano, J. T. Chiu, et al., “Simple and effective masked diffusion language models,” Advances in Neural Information Processing Systems, vol. 37, pp. 130136-130184, 2024, doi: 10.52202/079017-4135.
H. Zhang, K. Chen, X. Bai, Y. Xiang, and M. Zhang, “Paying more attention to source context: Mitigating unfaithful translations from large language model,” arXiv preprint arXiv:2406.07036, 2024.
L. Yang, H. Chen, Z. Li, X. Ding, and X. D. Wu, “Give us the facts: Enhancing large language models with knowledge graphs for fact-aware language modeling,” IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 7, pp. 3091-3110, 2024, doi: 10.1109/TKDE.2024.3360454.
M. Abulaish, M. Fazil, and M. J. Zaki, “Domain-specific keyword extraction using joint modeling of local and global contextual semantics,” ACM Transactions on Knowledge Discovery from Data (TKDD), vol. 16, no. 4, pp. 1-30, 2022, doi: 10.1145/3494560.
S. Kılıçarslan, K. Adem, and M. Çelik, “An overview of the activation functions used in deep learning algorithms,” Journal of New Results in Science, vol. 10, no. 3, pp. 75-88, 2021, doi: 10.54187/jnrs.1011739.
S. Abinaya, M. Deepak, and A. S. Alphonse, “Enhanced Image Captioning Using Bahdanau Attention Mechanism and Heuristic Beam Search Algorithm,” IEEE Access, vol. 12, pp. 78945-78956, 2024, doi: 10.1109/ACCESS.2024.3431091.
I. R. Visan, “Errors and Difficulties in Translating Maritime Terminology,” Translation Studies: Theory and Practice, vol. 3, no. 1, pp. 66-84, 2023, doi: 10.46991/TSTP/2023.3.1.066.
X. Wen and W. Li, “Time series prediction based on LSTM-attention-LSTM model,” IEEE access, vol. 11, pp. 48322-48331, 2023, doi: 10.1109/ACCESS.2023.3276628.
S. Lv, S. Lu, R. Wang, L. Yin, Z. Yin, S. A.AIQahtani, et al., “Enhancing Chinese Dialogue Generation with Word– Phrase Fusion Embedding and Sparse SoftMax Optimization,” Systems, vol. 12, no. 12, pp. 516-516, 2024, doi: 10.3390/systems12120516.
Y. Hao, Y. Liu, and L. Mou, “Teacher forcing recovers reward functions for text generation,” Advances in Neural Information Processing Systems, vol. 35, pp. 12594-12607, 2022, [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/hash/51ae7d9db3423ae96cd6afeb01529819-Abstract-Conference.html.
J. Lu and S. Steinerberger, “Neural collapse under cross-entropy loss,” Applied and Computational Harmonic Analysis, vol. 59, pp. 224-241, 2022, doi: 10.1016/j.acha.2021.12.011.
J. Choo, D. Kwon Y, J. Kim, J. Jae, A. Hottung, K. Tierney, et al., “Simulation-guided beam search for neural combinatorial optimization,” Advances in Neural Information Processing Systems, vol. 35, pp. 8760-8772, 2022, [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/hash/39b9b60f0d149eabd1fff2d7c7d5afc4-Abstract-Conference.html.
M. R. Islam, A. A. Lima, S. C. Das, M. F. Mridha, A. R. Prodeep, and Y. Watanobe, “A comprehensive survey on the process, methods, evaluation, and challenges of feature selection,” IEEE Access, vol. 10, pp. 99595-99632, 2022, doi: 10.1109/ACCESS.2022.3205618.