A Method for Alleviating Illusions in Code Generation Based on a Large Language Model Generated by Retrieval Enhancement

Main Article Content

K. Wei

Abstract

Large language models have achieved strong performance in code generation, but generated code may contain syntactically plausible yet semantically incorrect API calls, invalid imports, and inconsistent logic. To mitigate such code -generation hallucinations, this study proposes a retrieval-augmented generation framework integrating multi-source retrieval, generation constraints, post-verification, and feedback optimization. Code repositories, technical documents, and Q&A data are jointly modeled through semantic and structural retrieval, and retrieved fragments are represented as structured context units containing API descriptions, sample code, and constraints. During generation, token-level confidence assessment and semantic-consistency constraints are used to identify and control low-confidence fragments. After generation, abstract syntax tree analysis, type consistency checking, API constraint checking, and unittest feedback form a closed-loop generation–verification–correction process. A retrieval–generation collaborative optimization mechanism further updates retrieval weights according to confidence, consistency, and execution feedback. Experiments on code-generation benchmarks show that the method increases BLEU to 42.7, CodeBLEU to 46.3, and Pass@1 to 56.4, while reducing the hallucination rate to 9.8%. The framework provides a reliable method for semantic retrieval, software-engineering automation, and constrained intelligent code generation.

Downloads

Download data is not yet available.

Article Details

How to Cite
Wei, K. (2026). A Method for Alleviating Illusions in Code Generation Based on a Large Language Model Generated by Retrieval Enhancement. Advanced Electromagnetics, 15(3), 8519–8525. https://doi.org/10.7716/aem.v15i3.3976
Section
Research Articles

References

N. Chakraborty, M. Ornik, and K. Driggs-Campbell, “Hallucination detection in foundation models for decision-making: A flexible definition and review of the state of the art,” ACM Computing Surveys, vol. 57, no. 7, pp. 1-35, 2025, doi: 10.1145/3716846.

View Article

X. Yu, L. Liu, X. Hu, et al., “Fight fire with fire: How much can we trust chatgpt on source code-related tasks?” IEEE Transactions on Software Engineering, vol. 50, no. 12, pp. 3435-3453, 2024, doi: 10.1109/TSE.2024.3492204.

View Article

Y. Sun, D. Sheng, Z. Zhou, et al., “AI hallucination: towards a comprehensive classification of distorted information in artificial intelligence-generated content,” Humanities and Social Sciences Communications, vol. 11, no. 1, pp. 1-14, 2024, doi: 10.1057/s41599-024-03811-x.

View Article

X. Li, J. Jin, Y. Zhou, et al., “From matching to generation: A survey on generative information retrieval,” ACM Transactions on Information Systems, vol. 43, no. 3, pp. 1-62, 2025, doi: 10.1145/3722552.

View Article

Y. Lyu, Z. Li, S. Niu, et al., “Crud-rag: A comprehensive chinese benchmark for retrieval-augmented generation of large language models,” ACM Transactions on Information Systems, vol. 43, no. 2, pp. 1-32, 2025, doi: 10.1145/3701228.

View Article

L. Hu, Z. Liu, Z. Zhao, et al., “A survey of knowledge enhanced pre-trained language models,” IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 4, pp. 1413-1430, 2023, doi: 10.1109/TKDE.2023.3310002.

View Article

B. Madhumala R, B. Vineetha, R. Shree M, et al., “A semantic driven model for extraction of text using TF-IDF, SFLA and XGBoost,” International Journal of Information Technology, vol. 17, no. 6, pp. 3659-3664, 2025, doi: 10.1007/s41870-025-02549-2.

View Article

Z. Ji, N. Lee, R. Frieske, et al., “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, pp. 1-38, 2023, doi: 10.1145/3571730.

View Article