Research on BERT Model Compression and Acceleration Based on Knowledge Distillation and Dynamic Routing
Main Article Content
Abstract
The BERT model has achieved strong performance in natural language processing, but its large parameter scale and high computational cost limit deployment in resource-constrained and real-time scenarios. This study proposes a lightweight BERT compression and acceleration method that integrates knowledge distillation and dynamic routing. In the knowledge distillation stage, a multi-layer feature alignment and attention transfer mechanism is developed to transfer deep semantic knowledge from the teacher model to a compact student model. The strategy uses output-layer distillation, intermediate-layer feature alignment, and attention-map transfer to preserve semantic representation ability. In the dynamic routing stage, an adaptive inference structure based on early exit is constructed, enabling the model to adjust computational depth according to the complexity of different input samples and reduce redundant computation. A joint optimization strategy is used to balance final-output accuracy, early-exit classifier performance, and distillation effectiveness. Experimental results on multiple natural language processing tasks show that the proposed method substantially reduces parameter count, storage space, and inference latency while maintaining competitive performance relative to the original BERT model.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
M. Paganelli, F. Del Buono, A. Baraldi, et al., “Analyzing how BERT performs entity matching, ” Proceedings of the VLDB Endowment, vol. 15, no. 8, pp. 1726-1738, 2022, doi: 10.14778/3529337.3529356.
S. Kondra, V. Raghavan, and V. Kumar Adari, “Beyond Text: Exploring Multimodal BERT Models, ” International Journal of Research Publications in Engineering, Technology and Management (IJRPETM), vol. 8, no. 1, pp. 11764-11769, 2025, doi: 10.15662/IJRPETM.2025.0801003.
M. M. Danyal, S. S. Khan, M. Khan, et al., “Proposing sentiment analysis model based on BERT and XLNet for movie reviews, ” Multimedia Tools and Applications, vol. 83, no. 24, pp. 64315-64339, 2024, doi: 10.1007/s11042-024-18156-5.
I. Karabila, N. Darraz, A. El-Ansari, et al., “BERT-enhanced sentiment analysis for personalized e-commerce recommendations, ” Multimedia Tools and Applications, vol. 83, no. 19, pp. 56463-56488, 2024, doi: 10.1007/s11042-023-17689-5.
C. Özkurt, “Comparative analysis of state-of-the-art Q&A models: BERT, RoBERTa, DistilBERT, and ALBERT on squad v2 dataset, ” Chaos and Fractals, vol. 1, no. 1, pp. 19-30, 2024, doi: 10.69882/adba.chf.2024073.
A. C. Mazari, N. Boudoukhani, and A. Djeffal, “BERT-based ensemble learning for multi-aspect hate speech detection, ” Cluster Computing, vol. 27, no. 1, pp. 325-339, 2024, doi: 10.1007/s10586-022-03956-x.
A. Deshmukh and A. Raut, “Applying bert-based nlp for automated resume screening and candidate ranking, ” Annals of Data Science, vol. 12, no. 2, pp. 591-603, 2025, doi: 10.1007/s40745-024-00524-5.
M. Mujahid, K. Kanwal, F. Rustam, et al., “Arabic ChatGPT tweets classification using RoBERTa and BERT ensemble model, ” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 22, no. 8, pp. 1-23, 2023, doi: 10.1145/3605889.
M. Z. Zaki, “Revolutionising translation technology: A comparative study of variant transformer models–BERT, GPT and T5, ” Computer Science and Engineering–An International Journal, vol. 14, no. 3, pp. 15-27, 2024, doi: 10.5121/cseij.2024.14302.
M. N. Uddin, B. Li, Z. Ali, et al., “Software defect prediction employing BiLSTM and BERT-based semantic feature, ” Soft Computing, vol. 26, no. 16, pp. 7877-7891, 2022, doi: 10.1007/s00500-022-06830-5.
M. Gupta and P. Agrawal, “Compression of deep learning models for text: A survey, ” ACM Transactions on Knowledge Discovery from Data (TKDD), vol. 16, no. 4, pp. 1-55, 2022, doi: 10.1145/3487045.
E. C. Garrido-Merchan, R. Gozalo-Brizuela, and S. Gonzalez-Carvajal, “Comparing BERT against traditional machine learning models in text classification, ” Journal of Computational and Cognitive Engineering, vol. 2, no. 4, pp. 352-356, 2023, doi: 10.47852/bonviewJCCE3202838.
L. George and P. Sumathy, “An integrated clustering and BERT framework for improved topic modeling, ” International Journal of Information Technology, vol. 15, no. 4, pp. 2187-2195, 2023, doi: 10.1007/s41870-023-01268-w.
M. Bilal and A. A. Almazroi, “Effectiveness of Fine-tuned BERT Model in Classification of Helpful and Unhelpful Online Customer Reviews: M. Bilal, AA Almazroi, ” Electronic Commerce Research, vol. 23, no. 4, pp. 2737-2757, 2023, doi: 10.1007/s10660-022-09560-w.
M. Laurer, W. Van Atteveldt, A. Casas, et al., “Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and bert-nli, ” Political Analysis, vol. 32, no. 1, pp. 84-100, 2024, doi: 10.1017/pan.2023.20.