Kindergarten Safety Education Recommendation Based on Multimodal Fusion and ALBERT Lightweight Model
Main Article Content
Abstract
Current mainstream multimodal recommendation models for kindergarten safety education usually have large parameter scales and high computational overhead, which makes them difficult to deploy on resource-constrained edge devices. Similar deployment constraints also exist in engineering systems involving wireless sensing, electromagnetic environment monitoring, and antenna-assisted edge perception, where real-time inference and low resource consumption are required. To balance lightweight design and multimodal information integration, this paper proposes a kindergarten safety education recommendation framework based on multimodal fusion and an ALBERT lightweight model. The framework uses ALBERT-base-v2 with frozen embedding layers and fine-tunes only the last two Transformer encoder layers as the text encoder. A lightweight ResNet-18 and WaveNet are adopted for image and audio feature extraction, respectively. Trimodal semantic alignment is achieved through 128-dimensional linear projection, and a low-overhead gating mechanism is introduced to dynamically integrate modal features. Finally, a dual-tower architecture and Bayesian Personalized Ranking (BPR) loss are used to realize personalized recommendation. Experimental results show that the proposed model achieves an inference latency of 86.67 ms, a model size of 87.3 MB, an average HR@5 of 91.4%, and an average age-appropriateness score of 4.32. The results indicate that the model effectively balances deployment efficiency, multimodal representation capability, and adaptability to children’s development, providing a practical solution for intelligent kindergarten safety education systems and lightweight edge recommendation applications related to electromagnetic-aware sensing environments.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
L. S. PEK, R. W. M. MEE, W. Y. VON, M. R. ISMAIL, A. B. D. GHANI KHATIPAH, and T. S. T. SHAHDAN, “Be safe or be sorry: A scoping review on children’s safety and security awareness,” Quantum Journal of Social Sciences and Humanities, vol. 3, no. 5, pp. 14-25, 2022, doi: 10.55197/qjssh.v3i5.170.
C. Fosu-Ayarkwah, G. Gyeabour Fosu, and I. Awortwe, “Effects of teachers’ supervision on the safety of kindergarten pupils in the central region of Ghana,” Open Journal of Educational Research, vol. 2, no. 6, pp. 355-366, 2022, doi: 10.31586/ojer.2022.542.
N. Shaari, N. F. paiman, Y. ahmad, A. H. A. abd ghani, H. makhpol, A. M. radzi, et al., “Child Occupant Safety Program: A Study at Selected Kindergartens in Kajang and Putrajaya,” Journal of the Society of Automotive Engineers Malaysia, vol. 9, no. 1, pp. 23-28, 2025, doi: 10.56381/jsaem.v2i1.75.
Y. Yuan, Z. Li, and B. Zhao, “A survey of multimodal learning: Methods, applications, and future,” ACM Computing Surveys, vol. 57, no. 7, pp. 1-34, 2025, doi: 10.1145/3713070.
J. Moon, S. Yeo, Banihashem Sk, and O. Noroozi, “Using multimodal learning analytics as a formative assessment tool: Exploring collaborative dynamics in mathematics teacher education,” Journal of Computer Assisted Learning, vol. 40, no. 6, pp. 2753-2771, 2024, doi: 10.1111/jcal.13028.
A. Yusuf, N. M. Noor, and S. Bello, “Using multimodal learning analytics to model students’ learning behavior in animated programming classroom,” Education and Information Technologies, vol. 29, no. 6, pp. 6947-6990, 2024, doi: 10.1007/s10639-023-12079-8.
X. Wang, G. Chen, G. Qian, P. Gao, X. Y. Wei, and Y. Wang, “Large-scale multi-modal pre-trained models: A comprehensive survey,” Machine Intelligence Research, vol. 20, no. 4, pp. 447-482, 2023, doi: 10.1007/s11633-022-1410-8.
X. Han, Y. T. Wang, J. L. Feng, C. Deng, Z. H. Chen, Y. A. Huang, et al., “A survey of transformer-based multimodal pre-trained modals,” Neuro-computing, vol. 515, no. 1, pp. 89-106, 2023, doi: 10.1016/j.neucom.2022.09.136.
H. I. Liu, M. Galindo, H. Xie, L. K. Wong, H. H. Shuai, Y. H. Li, et al., “Lightweight deep learning for resource-constrained environments: A sur-vey,” ACM Computing Surveys, vol. 56, no. 10, pp. 1-42, 2024, doi: 10.1145/3657282.
M. Ponti, “Screen time and preschool children: Promoting health and development in a digital world,” Paediatrics & Child Health, vol. 28, no. 3, pp. 184-192, 2023, doi: 10.1093/pch/pxac125.
M. Oybarchin, “Psychological and pedagogical aspects of the development of visual activities of preschool children,” qo ’qon universiteti xabarnomasi, vol. 8, no. 1, pp. 101-105, 2023, doi: 10.54613/ku.v8i8.815.
O. V. Zashchirinskaia, “Nonverbal patterns of preschooler’s perception of visual images with the help of eye-tracker method usage,” Current psychology, vol. 40, no. 1, pp. 442-453, 2021, doi: 10.1007/s12144-018-9960-1.
G. De Mello, M. N. A. Ibrahim, N. Arumugam, M. S. Husin, N. H. O. Ma’mor, and S. Dharinee, “Nursery rhymes: Its effectiveness in teaching of English among pre-schoolers,” International Journal of Academic Research in Business and Social Sciences, vol. 12, no. 6, pp. 1914-1924, 2022, doi: 10.6007/IJARBSS/v12-i6/14124.
T. N. Fitria, “Using nursery rhymes in teaching English for young learners at childhood education,” Athena: Journal of Social, Culture and Society, vol. 1, no. 2, pp. 58-66, 2023, doi: 10.58905/athena.v1i2.28.
Y. Christina and P. Pujiarto, “The Effectiveness of Nursery Rhymes Media to Improve English Vocabulary and Confidence of Children (4-5 Years) in Tutor Time Kindergarten,” Journal of Education Research, vol. 4, no. 3, pp. 1326-1333, 2023, doi: 10.37985/jer.v4i3.406.
F. Hayat, “The effect of education using video animation on elementary school in hand washing skill,” Acitya: Journal of Teaching and Education, vol. 3, no. 1, pp. 44-53, 2021, doi: 10.30650/ajte.v3i1.2135.
M. Tuo and B. Long, “Construction and application of a human-computer collaborative multimodal practice teaching model for preschool education,” Computational Intelligence and Neuroscience, vol. 2022, no. 1, Art. no. 2973954, 2022, doi: 10.1155/2022/2973954.
E. Fundelius, T. Wade, A. Robbins, S. Wang, M. A. McConomy, and K. Fumero, “Universal design principles for multimodal representation in litera-cy activities for preschoolers,” Inclusive Practices, vol. 2, no. 1, pp. 13-21, 2023, doi: 10.1177/27324745221140380.
F. Ally, J. D. Pillay, and N. Govender, “Teaching and learning considerations during the COVID-19 pandemic: Supporting multimodal student learning preferences,” African Journal of Health Professions Education, vol. 14, no. 1, pp. 13-16, 2022, doi: 10.7196/AJHPE.2022.v14i1.1468.
Q. Si, T. S. Hodges, and J. M. Coleman, “Multimodal literacies classroom instruction for K-12 students: a review of research,” Literacy Research and Instruction, vol. 61, no. 3, pp. 276-297, 2022, doi: 10.1080/19388071.2021.2008555.
X. Ren, W. Yang, X. Jiang, G. Jin, and Y. Yu, “A deep learning framework for multimodal course recommendation based on LSTM+ attention,” Sustainability, vol. 14, no. 5, pp. 2907, 2022, doi: 10.3390/su14052907.
M. Ryu, G. Lee, and K. Lee, “Knowledge distillation for bert unsupervised domain adaptation,” Knowledge and Information Systems, vol. 64, no. 11, pp. 3113-3128, 2022, doi: 10.1007/s10115-022-01736-y.
I. Cho and U. Kang, “Pea-KD: Parameter-efficient and accurate Knowledge Distillation on BERT,” Plos One, vol. 17, no. 2, Art. no. e0263592, 2022, doi: 10.1371/journal.pone.0263592.
S. Safitri, M. Alii, and O. Mahmud, “Murottal Audio as a Medium for Memorizing the Qur’an in Super-Active Chil-dren,” Journal International Inspire Education Technology, vol. 1, no. 2, pp. 111-124, 2022, doi: 10.55849/jiiet.v1i2.87.
S. K. Uppada, P. Patel, and S. B, “An image and text-based multimodal model for detecting fake news in OSN’s,” Journal of Intelligent Information Systems, vol. 61, no. 2, pp. 367-393, 2023, doi: 10.1007/s10844-022-00764-y.
A. Chen, S. Ambrogio, P. Narayanan, A. Okazaki, C. Mackin, A. Fasoli, et al., “Demonstration of transformerbased ALBERT model on a 14nm analog AI inference chip,” Nature Communications, vol. 16, no. 1, pp. 8661, 2025, doi: 10.1038/s41467-025-63794-4.
M. S. Sidhu, N. A. A. Latib, and K. K. Sidhu, “MFCC in audio signal processing for voice disorder: a review,” Multimedia Tools and Applications, vol. 84, no. 10, pp. 8015-8035, 2025, doi: 10.1007/s11042-024-19253-1.
A. I. Siam, A. A. Elazm, N. A. El-Bahnasawy, G. M. El Banby, and F. E. Abd El-Samie, “PPG-based human identification using Mel-frequency cepstral coefficients and neural networks,” Multimedia Tools and Applications, vol. 80, no. 17, pp. 26001-26019, 2021, doi: 10.1007/s11042-021-10781-8.
S. Mehra, V. Ranga, and R. Agarwal, “Multimodal Integration of Mel Spectrograms and Text Transcripts for Enhanced Automatic Speech Recognition: Leveraging Extractive Transformer-Based Approaches and Late Fusion Strategies,” Computational Intelligence, vol. 40, no. 6, Art. no. e70012, 2024, doi: 10.1111/coin.70012.
U. R. Pol, P. S. Vadar, and T. T. Moharekar, “Hugging Face: Revolutionizing AI and NLP,” International Journal for Research in Applied Science and Engineering Technology, vol. 12, no. 8, pp. 1121-1124, 2024, doi: 10.22214/ijraset.2024.64023.
A. Ajibode, A. A. Bangash, F. R. Cogo, B. Adams, and A. E. Hassan, “Towards semantic versioning of open pre-trained language mod-el releases on hugging face,” Empirical Software Engineering, vol. 30, no. 3, pp. 1-63, 2025, doi: 10.1007/s10664-025-10631-3.
Z. Yin, E. Xing, and Z. Shen, “Squeeze, recover and relabel: Dataset condensation at imagenet scale from a new perspective,” Advances in Neural Information Processing Systems, vol. 36, no. 1, pp. 73582-73603, 2023, doi: 10.52202/075280-3219.
A. Ullah, H. Elahi, Z. Sun, A. Khatoon, and I. Ahmad, “Comparative analysis of AlexNet, ResNet18 and SqueezeNet with diverse modification and arduous implementation,” Arabian Journal for Science and Engineering, vol. 47, no. 2, pp. 2397-2417, 2022, doi: 10.1007/s13369-021-06182-6.
R. Helaly, S. Messaoud, S. Bouaafia, M. A. Hajjaji, and A. Mtibaa, “DTL-I-ResNet18: facial emotion recognition based on deep trans-fer learning and improved ResNet18,” Signal, Image and Video Processing, vol. 17, no. 6, pp. 2731-2744, 2023, doi: 10.1007/s11760-023-02490-6.
A. Paul, Z. Wu, K. Liu, and S. Gong, “Robust multi-objective visual bayesian personalized ranking for multimedia recommendation,” Applied Intelligence, vol. 52, no. 4, pp. 3499-3510, 2022, doi: 10.1007/s10489-021-02355-w.