Multimodal Data-Driven Intelligent Assessment System for Classroom Teaching Quality—Based on Deep Learning Methods and Practices
Main Article Content
Abstract
With the rapid advancement of intelligent sensing technologies and Electromagnetic Waves, Antennas and Propagation systems, efficient multimodal information fusion and temporal perception have become fundamental requirements for communication-oriented monitoring and intelligent decision-making. To address the challenges of heterogeneous data integration and subjective classroom assessment, this study proposes a Hierarchical Cross-modal Attention with Temporal Dynamics Perception Network (HCA-TDPNet) for multimodal teaching quality evaluation. The framework integrates Video Swin Transformer, Wav2Vec 2.0, and BERT to extract visual, acoustic, and semantic representations, while a multi-head cross-attention mechanism and a hierarchical temporal encoder composed of temporal convolutional networks and gated recurrent units enable deep cross-modal interaction and long-term temporal dependency modeling. The resulting multimodal representations are mapped to quantitative indicators of teaching objective achievement, teacher-student interaction quality, and student participation depth. Experimental validation based on public pre-training datasets and the SCB-Dataset demonstrates Pearson correlation coefficients of 0.893, 0.918, and 0.904 for the three evaluation dimensions, respectively, while reducing the assessment time of a single lesson to approximately 2 minutes, saving more than 95% compared with manual evaluation. Beyond educational applications, the proposed architecture provides an effective framework for heterogeneous information fusion, real-time multimodal perception, and adaptive semantic interaction, offering valuable engineering references for intelligent sensing, communication-oriented monitoring, and distributed information processing in Electromagnetic Waves, Antennas and Propagation systems.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
M. S. Far, “Artificial intelligence, innovative educational technologies, and transformation in classroom assessment,” Metaversalize, vol. 2, no. 2, pp. 122-131, 2025, doi: 10.22105/metaverse.v2i2.70.
M. Kovacheva, “From Big Data To Innovation: Examining The Impact Of Digitalization And Artificial Intelligence On Education,” Nar-o dnoctopancki arxiv, vol. (1), pp. 94-110, 2025, doi: 10.58861/tae.ea-nsa.2025.1.05.en.
N. G. Bepari, R. Barua, and F. Rabbi, “Quality Assurance In Education: Innovative Approaches For Effective Classroom Management And Student Engagement,” Journal Of Creative Writing, vol. 8, no. 3, pp. 1-22, 2024, doi: 10.1177/23476311221143231.
Y. Gao, “Deep learning-based strategies for evaluating and enhancing university teaching quality,” Computers and Education: Artificial Intelligence,v ol. 8, no. 1, Art. no. 100362, 2025, doi: 10.1016/j.caeai.2025.100362.
A. Barua, M. U. Ahmed, and S. Begum, “A systematic literature review on multimodal machine learning: Applications, challenges, gaps and future directions,” Ieee access, vol. 11, no. 1, pp. 14804-14831, 2023, doi: 10.1109/ACCESS.2023.3243854.
S. Song, X. Li, S. Li, et al., “How to bridge the gap between modalities: Survey on multimodal large language model,” IEEE Transactions on Knowledge and Data Engineering, vol. 37, no. 9, pp. 5311-5329, 2025, doi: 10.1109/TKDE.2025.3527978.
W. Zhang, S. Gan, H. He, et al., “Effects of Teaching Behaviors on the Effectiveness of Classroom Learning Time,” International Journal of Emerging Technologies in Learning, vol. 17, no. 9, pp. 214-227, 2022, doi: 10.3991/ijet.v17i09.30943.
M. Krichen and A. Mihoub, “Long short-term memory networks: A comprehensive survey,” AI, vol. 6, no. 9, pp. 215, 2025, doi: 10.3390/ai6090215.
C. H. Heng, M. Toyoura, C. S. Leow, et al., “Analysis of Classroom Processes Based on Deep Learning With Video and Audio Features,” IEEE Access, vol. 12, no. 1, pp. 110705-110712, 2024, doi: 10.1109/ACCESS.2024.3434742.
N. I. Mohd Talib, N. A. Abd Majid, and S. Sahran, “Identification of student behavioral patterns in higher education using Kmeans clustering and support vector machine,” Applied Sciences, vol. 13, no. 5, pp. 3267, 2023, doi: 10.3390/app13053267.
T. Xu, W. Deng, S. Zhang, et al., “Research on recognition and analysis of teacher-student behavior based on a blended synchronous classroom,” Applied Sciences, vol. 13, no. 6, pp. 3432, 2023, doi: 10.3390/app13063432.
Y. Kumar, A. Koul, and C. Singh, “A deep learning approaches in text-to-speech system: a systematic review and recent research perspective,” Multimedia Tools and Applications, vol. 82, no. 10, pp. 15171-15197, 2023, doi: 10.1007/s11042-022-13943-4.
F. Zhao, C. Zhang, and B. Geng, “Deep multimodal data fusion,” ACM computing surveys, vol. 56, no. 9, pp. 1-36, 2024, doi: 10.1145/3649447.
T. Gao, Q. Zhang, T. Chen, et al., “Spatial-temporal sequence attention based efficient transformer for video snow removal,” Big Data Mining and Analytics, vol. 8, no. 3, pp. 551-562, 2025, doi: 10.26599/BDMA.2024.9020061.
Z. C. Liu, L. Chen, Y. J. Hu, et al., “Pe-wav2vec: A prosody-enhanced speech model for self-supervised prosody learning in tts,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, no. 1, pp. 4199-4210, 2024, doi: 10.1109/TASLP.2024.3449148.
A. Mohamed, H. Lee, L. Borgholt, et al., “Self-supervised speech representation learning: A review,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1179-1210, 2022, doi: 10.1109/JSTSP.2022.3207050.
M. Cao, F. Xu, H. Jia, et al., “A multiple interpolation algorithm to improve resampling accuracy in data triggers,” Electronics, vol. 12, no. 6, pp. 1291, 2023, doi: 10.3390/electronics12061291.
A. S. Parihar, D. Varshney, K. Pandya, et al., “A comprehensive survey on video frame interpolation techniques,” The Visual Computer, vol. 38, no. 1, pp. 295-319, 2022, doi: 10.1007/s00371-020-02016-y.
A. Fatwanto, F. Zamakhsyari, and R. Ndungi, “A systematic literature review of BERT-based models for natural language processing tasks,” JURNAL INFOTEL, vol. 16, no. 4, pp. 713-728, 2024, doi: 10.20895/infotel.v16i4.1206.
A. C. Mazari, N. Boudoukhani, and A. Djeffal, “BERT-based ensemble learning for multi-aspect hate speech detection,” Cluster Computing, vol. 27, no. 1, pp. 325-339, 2024, doi: 10.1007/s10586-022-03956-x.
J. Wu, X. Wang, X. Gao, et al., “On the effectiveness of sampled softmax loss for item recommendation,” ACM Transactions on Information Systems, vol. 42, no. 4, pp. 1-26, 2024, doi: 10.1145/3637061.
H. Kang, T. Lv, C. Yang, et al., “Multihead-res-se residual network with attention for human activity recognition,” Electronics, vol. 13, no. 17, pp. 3407, 2024, doi: 10.3390/electronics13173407.
C. Cao, J. Huang, M. Wu, et al., “A multivariate time series prediction method based on convolution-residual gated recurrent neural network and double-layer attention,” Electronics, vol. 13, no. 14, pp. 2834, 2024, doi: 10.3390/electronics13142834.
X. Yuan, X. Shen, S. Mehta, et al., “Structure injected weight normalization for training deep networks,” Multimedia Systems, vol. 28, no. 2, pp. 433-444, 2022, doi: 10.1007/s00530-021-00793-7.
L. Huang, J. Qin, Y. Zhou, et al., “Normalization techniques in training dnns: Methodology, analysis and application,” IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 8, pp. 10173-10196, 2023, doi: 10.1109/TPAMI.2023.3250241.
M. Kohler, J. Cho, and A. Krzyzak,˙ “On the rate of convergence of an over-parametrized deep neural network regression estimate with ReLU activation function learned by gradient descent,” Electronic Journal of Statistics, vol. 19, no. 2, pp. 4829-4866, 2025, doi: 10.1214/25-EJS2444.
H. Peng, B. Jiang, Z. Mao, et al., “Local enhancing transformer with temporal convolutional attention mechanism for bearings remaining useful life prediction,” IEEE Transactions on Instrumentation and Measurement, vol. 72, no. 1, pp. 1-12, 2023, doi: 10.1109/TIM.2023.3291787.
M. Lu and S. S. Shin, “Intelligent Temperature Control System Utilizing Gated Recurrent Units and the Internet of Things,” Journal of the Digital Content Society, vol. 25, no. 11, pp. 3419-3429, 2024, doi: 10.9728/dcs.2024.25.11.3419.
J. Sun, Z. Wei, and X. Liu, “GRU-based model-free adaptive control for industrial processes,” Neural Computing and Applications, vol. 35, no. 24, pp. 17701-17715, 2023, doi: 10.1007/s00521-023-08652-4.
Z. Pan, Z. Gu, X. Jiang, et al., “A modular approximation methodology for efficient fixed-point hardware implementation of the sigmoid function,” IEEE Transactions on Industrial Electronics, vol. 69, no. 10, pp. 10694-10703, 2022, doi: 10.1109/TIE.2022.3146573.
C. Martinelli, A. Coraddu, and A. Cammarano, “Approximating piecewise nonlinearities in dynamic systems with sigmoid functions: advantages and limitations,” Nonlinear Dynamics, vol. 111, no. 9, pp. 8545-8569, 2023, doi: 10.1007/s11071-023-08293-1.
W. Ali, S. Vascon, T. Stadelmann, et al., “Hierarchical glocal attention pooling for graph classification,” Pattern Recognition Letters, vol. 186, no. 1, pp. 71-77, 2024, doi: 10.1016/j.patrec.2024.09.009.
K. Ma, C. Tang, W. Zhang, et al., “DC-CNN: Dual-channel Convolutional Neural Networks with attention-pooling for fake news detection,” Applied Intelligence, vol. 53, no. 7, pp. 8354-8369, 2023, doi: 10.1007/s10489-022-03910-9.
Q. Jiang, L. Zhu, C. Shu, et al., “An efficient multilayer RBF neural network and its application to regression problems,” Neural computing and Applications, vol. 34, no. 6, pp. 4133-4150, 2022, doi: 10.1007/s00521-021-06373-0.
M. N. B. Adnan, W. M. A. W. Ahmad, N. A. Rahman, et al., “A robust hybrid methodology between applied linear regression model (alrm) and multilayer perceptron (mlp),” Bangladesh Journal of Medical Science, vol. 22, no. 1, pp. 38-46, 2023, doi: 10.3329/bjms.v22i1.61850.