Automatic Assessment Model for Chinese-English Scientific Translation Quality Based on Contrastive Learning
Main Article Content
Abstract
Assessing the quality of scientific literature translation remains challenging because of strong subjectivity, dense domain-specific terminology, and the limited availability of standardized reference translations. These issues are particularly relevant for the international dissemination of research in advanced electromagnetic engineering, where precise multilingual communication supports the reliable exchange of knowledge on electromagnetic waves, antennas, and propagation technologies. This paper proposes a Contrastive Learning-based Chinese-English Scientific Translation Quality Evaluation model (C-TQE). By constructing multi-level positive and negative sample pairs, the model learns the relative ordinal relationships of translation quality within a shared representation space. A dual-encoder architecture encodes source sentences and candidate translations through a shared pre-trained language model, while a contrastive loss function draws high-quality translations closer to the source representation and separates low-quality ones. To address the characteristics of scientific texts, a term-aware negative sampling strategy exploits domain dictionaries and syntactic structures to generate semantically similar but terminologically incorrect examples. Experiments on 11, 238 human-annotated instances from the WMT20–22 Chinese-English scientific translation tasks show that C-TQE achieves a Kendall’s tau correlation coefficient of 0.564 with human judgments, outperforming COMET (0.512) and BLEURT (0.497). Ablation studies confirm the effectiveness of term-aware negative sampling and the contrastive learning objective, while diagnostic analysis demonstrates high consistency in evaluating terminological accuracy and syntactic structures. The proposed framework provides an effective solution for large-scale scientific translation quality assessment and facilitates the accurate international communication of multidisciplinary engineering research, including electromagnetic and antenna-related studies.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
M. Karpinska, N. Raj, K. Thai, et al., “DEMETR: Diagnosing Evaluation Metrics for Translation,” Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9540-9561, 2022, doi: 10.18653/v1/2022.emnlp-main.649.
S. Gu and Y. Feng, “Research on Continuous Learning of Multilingual Neural Machine Translation Based on Parameter Allocation,” Journal of Chinese Information Processing, vol. 39, no. 4, pp. 77-84, 2025, doi: 10.3969/j.issn.1003-0077.2025.04.007.
X. Qu, G. Liu, and L. Li, “Multi-scale Joint Learning with Negative Sample Mining for Non-autoregressive Machine Translation,” Knowledge-Based Systems, vol. 322, Art. no. 113610, 2025, doi: 10.1016/j.knosys.2025.113610.
M. Zufle, V. Zouhar, and T. A. Dinh, “COMET-poly: Machine Translation Metric Grounded in Other Candidates,” Proceedings of the Tenth Conference on Machine Translation. Association for Computational Linguistics, pp. 887-904, 2025, doi: 10.18653/v1/2025.wmt-1.63.
J. Wang, “Research and Construction of Automatic Scoring Models for Chinese Students Chinese-to-English Translation Based on Different Styles,” Beijing: Foreign Language Teaching and Research Press, ISBN: 9787521358155, 2024.
M. Chopra, L. Sparrenberg, and R. Sifa, “SynCED-EnDe 2025: A Synthetic and Curated English-German Dataset for Critical Error Detection in Machine Translation,” arXiv preprint. arXiv:2510.05144, 2025.
T. Gao, X. Yao, and D. Chen, “SimCSE: Simple Contrastive Learning of Sentence Embeddings,” Conference on Empirical Methods in Natural Language Processing, pp. 6894-6910, 2021, doi: 10.18653/v1/2021.emnlp-main.552.
Evaluating Terminology Translation in Machine Translation Systems via Metamorphic Testing, “2024 IEEE International Conference on Software Testing, Verification and Validation (ICST),” IEEE, pp. 1-12, 2024.
Tongyi Lab, “WMT English-to-Chinese Machine Translation News Test Set,” ModelScope, 2022, [Online]. Available: http://www.modelscope.cn/datasets/damo/WMT-English-to-Chinese-Machine-Translation-newstest.git.
Q. Wang, “Evaluating Uighur literary translation: A comparative study of ChatGPT, Google Translate, and Bing Translator,” PLoS One, vol. 20, no. 10, Art. no. e0335261, 2025, doi: 10.1371/journal.pone.0335261.
X. Zheng, H. Chen, Y. Ma, et al., “Neural Machine Translation Model Fusing Dependency Syntax and LSTM,” Journal of Harbin University of Science and Technology, vol. 28, no. 3, pp. 20-27, 2023, doi: 10.15938/j.jhust.2023.03.003.
R. Rei, C. Stewart, A. C. Farinha, et al., “Comet: A Neural Framework for MT Evaluation,” Conference on Empirical Methods in Natural Language Processing, pp. 2685-2702, 2020, doi: 10.18653/v1/2020.emnlp-main.213.
T. Sellam, D. Das, and A. P. Parikh, “BLEURT: Learning Robust Metrics for Text Generation,” 58th Annual Meeting of the Association for Computational Linguistics, pp. 7881-7892, 2020, doi: 10.18653/v1/2020.acl-main.704.
T. Ranasinghe, C. Orasan, and R. Mitkov, “TransQuest: Translation Quality Estimation with Cross-lingual Transformers,” 28th International Conference on Computational Linguistics, pp. 5070-5081, 2020, doi: 10.18653/v1/2020.coling-main.445.
Y. Zheng and Y. Liu, “A Review of Machine Translation Quality Research,” English Square, vol. (23), pp. 40-43, 2025, doi: 10.3969/j.issn.1009-6167.2025.23.011.
N. Lin, “Research on Key Technologies of Multilingual Learning Based on Pre-trained Models,” Guangdong University of Technology, 2024, doi: 10.27029/d.cnki.ggdgu.2024.000095.