A Software Code Defect Detection Method Based on the Fusion of Transformer and Graph Neural Networks
Main Article Content
Abstract
Software defect detection plays a key role in improving system reliability, security, and maintainability. This study proposes a hybrid Transformer–GNN framework to jointly model code semantics and program structure. Source files are normalized, segmented at the function level, and converted into token sequences, while multi-relation graphs are generated to describe syntax, control flow, and data dependencies. The Transformer branch captures long-range contextual information from code tokens, whereas the GNN branch learns structural interactions among statements, execution paths, and variables. Their representations are adaptively integrated through a gated fusion module for defect classification. Experiments show that the proposed method outperforms conventional classifiers, sequence-oriented neural networks, and standalone Transformer or GNN models. Ablation results also verify the importance of semantic encoding, structural learning, data-flow information, and gated fusion. Overall, combining lexical context with graph-based program relations provides a more accurate, robust, and interpretable solution for software defect identification.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
N. S. Harzevili, A. B. Belle, J. Wang, et al., "A systematic literature review on automated software vulnerability detection using machine learning," ACM Comput. Surv., vol. 57, no. 3, pp. 55:1–55:36, 2024.
Y. Zheng, S. Pujar, B. Lewis, et al., "D2A: A dataset built for AI-based vulnerability detection methods using differential analysis," in Proc. 43rd IEEE/ACM Int. Conf. Softw. Eng.: Softw. Eng. Pract., Piscataway, NJ, USA: IEEE, 2021, pp. 111–120.
D. Guo, S. Ren, S. Lu, et al., "GraphCodeBERT: Pre-training code representations with data flow," in Proc. Int. Conf. Learn. Representations (ICLR), 2021.
W. U. Ahmad, S. Chakraborty, B. Ray, et al., "Unified pre-training for program understanding and generation," in Proc. 2021 Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Lang. Technol., Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 2655–2668.
Y. Wang, W. Wang, S. Joty, et al., "CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation," in Proc. 2021 Conf. Empirical Methods in Natural Language Processing (EMNLP), Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 8696–8708.
V. A. Nguyen, D. Q. Nguyen, V. Nguyen, et al., "ReGVD: Revisiting graph neural networks for vulnerability detection," in Proc. 44th Int. Conf. Softw. Eng. Companion, New York, NY, USA: ACM, 2022, pp. 51–55.
M. Fu and C. Tantithamthavorn, "LineVul: A transformer-based line-level vulnerability prediction," in Proc. 19th Int. Conf. Mining Softw. Repositories, New York, NY, USA: ACM, 2022, pp. 608–620.
H. Hin, A. Kan, H. Chen, et al., "LineVD: Statement-level vulnerability detection using graph neural networks," in Proc. 19th Int. Conf. Mining Softw. Repositories, New York, NY, USA: ACM, 2022, pp. 596–607.
H. Hanif and S. Maffeis, "VulBERTa: Simplified source code pre-training for vulnerability detection," in Proc. 2022 Int. Joint Conf. Neural Networks (IJCNN), Piscataway, NJ, USA: IEEE, 2022, pp. 1–8.
D. Guo, S. Lu, N. Duan, et al., "UniXcoder: Unified cross-modal pre-training for code representation," in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2022, pp. 7212–7225.
N. T. Islam, G. De La Torre Parra, D. Manuel, et al., "An unbiased transformer source code learning with semantic vulnerability graph," in Proc. 2023 IEEE 8th Eur. Symp. Security Privacy, Piscataway, NJ, USA: IEEE, 2023, pp. 144–159.
Y. Chen, Z. Ding, L. Alowain, et al., "DiverseVul: A new vulnerable source code dataset for deep learning based vulnerability detection," in Proc. 26th Int. Symp. Research Attacks, Intrusions Defense, New York, NY, USA: ACM, 2023, pp. 654–668.
F. Qiu, Z. Liu, X. Hu, et al., "Vulnerability detection via multiple-graph-based code representation," IEEE Trans. Softw. Eng., vol. 50, no. 8, pp. 2178–2199, 2024.
R. Liu, Y. Wang, H. Xu, et al., "Source code vulnerability detection: Combining code language models and code property graphs," arXiv preprint arXiv:2404.14719, 2024.
A. Lekssays, H. Mouhcine, K. Tran, et al., "LLMxCPG: Context-aware vulnerability detection through code property graph-guided large language models," in Proc. 34th USENIX Security Symp., Berkeley, CA, USA: USENIX Association, 2025.