Construction of Chinese-Japanese-English Translation Database Based on Cross-Lingual Transformer-Based Models

Main Article Content

X. D. Wang
D. L. Wu

Abstract

Accurate semantic alignment and comprehensive multilingual coverage remain major challenges in constructing Chinese–Japanese–English translation databases for technical knowledge sharing and engineering information exchange. This study proposes a cross-lingual translation database construction framework based on Transformerbased neural networks and hierarchical contrastive representation learning. Using InfoXLM as the shared semantic encoder, dependency-syntax perturbation and cross-lingual lexical substitution are employed to generate codeswitching hard negative samples, while a three-way triplet contrastive optimization strategy jointly constrains Chinese, Japanese, and English semantic spaces to improve representation consistency. The fine-tuned model is further integrated with semantic similarity filtering and bidirectional consistency verification to construct a large-scale trilingual translation database from multilingual candidate corpora. Experimental results demonstrate an alignment precision of 94.7%, a recall of 92.3%, an F1-score of 93.5%, and a translation coverage of 89.8%, significantly outperforming conventional multilingual alignment approaches. The proposed framework provides an efficient solution for multilingual engineering knowledge organization, technical document retrieval, and intelligent information processing, offering valuable support for cross-language electromagnetic engineering documentation, antenna technology resources, and multilingual communication systems requiring robust semantic alignment and signal-aware information management.

Downloads

Download data is not yet available.

Article Details

How to Cite
Wang, X. D., & Wu, D. L. (2026). Construction of Chinese-Japanese-English Translation Database Based on Cross-Lingual Transformer-Based Models. Advanced Electromagnetics, 15(3), 4124–4137. https://doi.org/10.7716/aem.v15i3.3477
Section
Research Articles

References

C. Mi and S. Xie, “Language relatedness evaluation for multilingual neural machine translation,” Neurocomputing, vol. 570, no. 1, Art. no. 127115, 2024, doi: 10.1016/j.neucom.2023.127115.

View Article

J. Luo, C. Cherry, and G. Foster, “To diverge or not to diverge: A morphosyntactic perspective on machine translation vs human translation,” Transactions of the Association for Computational Linguistics, vol. 12, no. 1, pp. 355-371, 2024, doi: 10.1162/tacl_a_00645.

View Article

S. Ranathunga, E. S. A. Lee, M. Prifti Skenduli, et al., “Neural machine translation for low-resource languages: A survey,” ACM Computing Surveys, vol. 55, no. 11, pp. 1-37, 2023, doi: 10.1145/3567592.

View Article

F. Li, C. Chi, H. Yan, et al., “STA: An efficient data augmentation method for low-resource neural machine translation,” Journal of Intelligent & Fuzzy Systems, vol. 45, no. 1, pp. 121-132, 2023, doi: 10.3233/jifs-230682.

View Article

A. Gui and H. Xiao, “Multi-level multilingual semantic alignment for zero-shot cross-lingual transfer learning,” Neural Networks, vol. 173, no. 1, pp. 106217-106231, 2024, doi: 10.1016/j.neunet.2024.106217.

View Article

Z. Mao, C. Chu, and S. Kurohashi, “Linguistically driven multi-task pre-training for low-resource neural machine translation,” Transactions on Asian and Low-Resource Language Information Processing, vol. 21, no. 4, pp. 1-29, 2022, doi: 10.1145/3491065.

View Article

H. Yanaka and K. Mineshima, “Compositional evaluation on Japanese textual entailment and similarity,” Transactions of the Association for Computational Linguistics, vol. 10, no. 1, pp. 1266-1284, 2022, doi: 10.1162/tacl_a_00518.

View Article

J. Zhang, Y. Tian, J. Mao, et al., “WCC-JC 2.0: A web-crawled and manually aligned parallel corpus for Japanese-Chinese neural machine translation,” Electronics, vol. 12, no. 5, pp. 1140-1157, 2023, doi: 10.3390/electronics12051140.

View Article

A. Fernando, S. Ranathunga, D. Sachintha, et al., “Exploiting bilingual lexicons to improve multilingual embedding-based document and sentence alignment for low-resource languages,” Knowledge and Information Systems, vol. 65, no. 2, pp. 571-612, 2023, doi: 10.1007/s10115-022-01761-x.

View Article

S. Zhu, S. Gu, S. Li, et al., “Mining parallel sentences from internet with multi-view knowledge distillation for low-resource language pairs,” Knowledge and Information Systems, vol. 66, no. 1, pp. 187-209, 2024, doi: 10.1007/s10115-023-01925-3.

View Article

Z. Li, J. Ma, and F. Ren, “Exploiting Multi-Level Data Uncertainty for Japanese-Chinese Neural Machine Translation,” IEICE Transactions on Information and Systems, vol. 108, no. 5, pp. 440-443, 2024, doi: 10.1587/transinf.2024edl8059.

View Article

Z. Man, Y. Zhang, Y. Li, et al., “An ensemble strategy with gradient conflict for multi-domain neural machine translation,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 2, pp. 1-22, 2024, doi: 10.1145/3638248.

View Article

J. Zhang, K. Su, H. Li, et al., “Neural machine translation for low-resource languages from a chinese-centric perspective: A survey,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 6, pp. 1-60, 2024, doi: 10.1145/3665244.

View Article

J. Zhang, K. Su, Y. Tian, et al., “WCC-EC 2.0: enhancing neural machine translation with a 1.6 M+ web-crawled English-Chinese parallel corpus,” Electronics, vol. 13, no. 7, pp. 1381-1392, 2024, doi: 10.3390/electronics13071381.

View Article

R. Choenni and E. Shutova, “Investigating language relationships in multilingual sentence encoders through the lens of linguistic typology,” Computational Linguistics, vol. 48, no. 3, pp. 635-672, 2022, doi: 10.1162/coli_a_00444.

View Article

Z. Mao, C. Chu, and S. Kurohashi, “Ems: Efficient and effective massively multilingual sentence embedding learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, no. 1, pp. 2841-2856, 2024, doi: 10.1109/TASLP.2024.3402064.

View Article

V. Mathur, T. Dadu, and S. Aggarwal, “Evaluating neural networks’ ability to generalize against adversarial attacks in crosslingual settings,” Applied Sciences, vol. 14, no. 13, pp. 5440-5463, 2024, doi: 10.3390/app14135440.

View Article

G. Tang, O. Yousuf, and Z. Jin, “Improving BERTScore for machine translation evaluation through contrastive learning,” IEEE access, vol. 12, no. 1, pp. 77739-77749, 2024, doi: 10.1109/access.2024.3406993.

View Article

L. Ding, L. Wang, and S. Liu, “Recurrent graph encoder for syntaxaware neural machine translation,” International Journal of Machine Learning and Cybernetics, vol. 14, no. 4, pp. 1053-1062, 2023, doi: 10.1007/s13042-022-01682-9.

View Article

F. Azadi, H. Faili, and M. J. Dousti, “Mismatching-aware unsupervised translation quality estimation for low-resource languages,” Languages Resources and Evaluation, vol. 58, no. 4, pp. 1207-1231, 2024, doi: 10.1007/s10579-024-09727-x.

View Article

B. Xing and I. W. Tsang, “HC^{2}2L: Hybrid and Cooperative Contrastive Learning for Cross-Lingual Spoken Language Understanding,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 8094-8105, 2024, doi: 10.1109/TPAMI.2024.3402746.

View Article

A. Banerjee, B. K. Singh, V. Kumar, et al., “Research Challenges and Future Directions in Transformer-Based Neural Machine Translation,” Expert Systems with Applications, vol. 1, no. 1, Art. no. 131062, 2026, doi: 10.1016/j.eswa.2025.131062.

View Article

I. Sel and D. Hanbay, “Efficient adaptation: enhancing multilingual models for low-resource language translation,” Mathematics, vol. 12, no. 19, pp. 3149-3156, 2024, doi: 10.3390/math12193149.

View Article

G. Tolegen, A. Toleu, and R. Mussabayev, “Contrastive learning for morphological disambiguation using large language models in low-resource settings,” Applied Sciences, vol. 14, no. 21, pp. 9992, 2024, doi: 10.3390/app14219992.

View Article

X. Chen, Y. Yang, and H. Hu, “Automatic evaluation of English translation based on multi-granularity interaction fusion,” Neural Processing Letters, vol. 57, no. 1, pp. 19-26, 2025, doi: 10.1007/s11063-025-11716-2.

View Article

V. K. Pant, R. Sharma, and S. Kundu, “An overview of stemming and lemmatization techniques,” Advances in networks, intelligence and computing, vol. 1, no. 1, pp. 308-321, 2024, doi: 10.1201/9781003430421-31.

View Article

J. Nivre, A. Basirat, L. Dürlich, et al., “Nucleus composition in transition-based dependency parsing,” Computational Linguistics, vol. 48, no. 4, pp. 849-886, 2022, doi: 10.1162/coli_a_00450.

View Article

S. Duan, H. Zhao, and D. Zhang, “Syntax-aware data augmentation for neural machine translation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, no. 1, pp. 2988-2999, 2023, doi: 10.1109/taslp.2023.3301214.

View Article

L. Li, A. Zhang, and M. X. Luo, “DSISA: A new neural machine translation combining dependency weight and neighbors,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 2, pp. 1-16, 2024, doi: 10.1145/3638762.

View Article

X. Liang, Z. Gu, Y. Xie, L. Wang, and Z. Tian, “MUSEDA: multilingual unsupervised and supervised embedding for domain adaption,” Knowledge-Based Systems, vol. 273, no. 1, pp. 110560-110578, 2023, doi: 10.1016/j.knosys.2023.110560.

View Article

Q. Tan, X. Song, G. Ye, et al., “An effective negative sampling approach for contrastive learning of sentence embedding,” Machine Learning, vol. 112, no. 12, pp. 4837-4861, 2023, doi: 10.1007/s10994-023-06408-8.

View Article

B. Kairatuly and A. Shomanov, “Balancing speed and performance with layer freezing strategies for transformer models,” Scientific Journal of Astana IT University, vol. 22, no. 1, pp. 153-162, 2025, doi: 10.37943/22oxky5402.

View Article

T. Hwang, H. Seo, J. Jung, et al., “Exploring Selective Layer Freezing Strategies in Transformer Fine-Tuning: NLI Classifiers with Sub-3B Parameter Models,” Applied Sciences, vol. 15, no. 19, pp. 10434-10453, 2025, doi: 10.3390/app151910434.

View Article

F. F. Ijebu, Y. Liu, C. Sun, et al., “Soft cosine and extended cosine adaptation for pre-trained language model semantic vector analysis,” Applied Soft Computing, vol. 169, no. 1, Art. no. 112551, 2025, doi: 10.1016/j.asoc.2024.112551.

View Article

F. R. Zagatti, G. Yuuji Shimizu, D. Lucredio, et al., “Investigating the Relationship Between Text Vectorization Cosine Similarity and Classification Performance,” IEEE Access, vol. 13, no. 1, pp. 137348-137363, 2025, doi: 10.1109/access.2025.3595423.

View Article

Similar Articles

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > >> 

You may also start an advanced similarity search for this article.