Music Teaching Resource Recommendation with Multimodal Transformers
Main Article Content
Abstract
Multimodal semantic fragmentation and inaccurate user-intent modeling reduce the effectiveness of music teaching resource recommendation. To address these problems, this paper proposes MMT-MERec, a Multimodal Transformer for Music Educational Recommendation framework that integrates a multimodal Transformer with an educational knowledge graph. The method extracts 768-dimensional semantic vectors from audio, sheet music, video, and text through four pre-trained models: AST, MusicBERT, VideoMAE, and Sentence-BERT. A six-layer Cross-Modal Transformer encoder then performs fine-grained semantic fusion, using audio as a query to align the other modalities. A RotatE-embedded Skill Knowledge Graph containing 142 nodes constrains the recommendation loss to maintain teaching logic. A Skill-Aware Contrastive Learning objective based on user dynamics, including error rate and session duration, is used to optimize resource ranking. On an 8,200-sample dataset, MMT-MERec significantly outperforms the M3oE baseline, achieving HR@5 of 0.603 and NDCG@10 of 0.537. In cold-start scenarios, Cold-Start HR@5 reaches 0.401. These results show that the model jointly optimizes multimodal semantic alignment and educational knowledge constraints, improving recommendation relevance and teaching effectiveness.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
P. Fang, “Optimization of music teaching in colleges and universities based on multimedia technology,” Advances in Educational Technology and Psychology, vol. 5, no. 5, pp. 47-57, 2021, doi: 10.23977/aetp.2021.55009.
D. A. Camlin and T. Lisboa, “The digital ‘turn’ in music education,” Music Education Research, vol. 23, no. 2, pp. 129 - 138, 2021, doi: 10.1080/14613808.2021.1908792.
Y. Shi and X. Yang, “A personalized matching system for management teaching resources based on collaborative filtering algorithm,” International Journal of Emerging Technologies in Learning (iJET), vol. 15, no. 13, pp. 207-220, 2020, doi: 10.3991/ijet.v15i13.15353.
Y. Wu and P. Chen, “Music recommendation model based on user long-term and short-term preferences and music emotional attention,” Journal of Guangdong University of Technology, vol. 40, no. 4, pp. 37-44, 2023, doi: 10.12052/gdutxb.220009.
H. Chen, G. Wu, J. Li, J. Wang, and Tao Hong, “Research progress of deep learning recommendation based on attention mechanism,” Computer Engineering & Science, vol. 43, no. 02, pp. 370, 2021, doi: 10.3969/j.issn.1007-130X.2021.02.023.
L. Liu, M. Kong, C. Cao, Z. Shu, K. Liu, X. Li, et al., “Personalized music recommendation algorithm based on machine learning,” Multimedia Systems, vol. 31, no. 2, pp. 166, 2025, doi: 10.1007/s00530-025-01749-x.
N. Huang, R. Hu, M. Xiong, X. Peng, H. Ding, X. Jia, et al., “Multi-scale interest dynamic hierarchical transformer for sequential recommendation,” Neural Computing and Applications, vol. 34, no. 19, pp. 16643-16654, 2022, doi: 10.1007/s00521-022-07281-7.
A. Paul, Z. Wu, K. Liu, and S. Gong, “Robust multi-objective visual bayesian personalized ranking for multimedia recommendation,” Applied Intelligence, vol. 52, no. 4, pp. 3499-3510, 2022, doi: 10.1007/s10489-021-02355-w.
G. Li, J. Zhuo, G. Xu, C. Li, G. Wu, and H. Zhang, “Correlation Visual Adversarial Bayesian Personalized Ranking Recommendation Model,” Advanced Engineering Science/Gongcheng Kexue Yu Jishu, vol. 54, no. 3, pp. 230, 2022, doi: 10.15961/j.jsuese.202100569.
K. Liu, F. Xue, S. Li, S. Sang, and R. Hong, “Multimodal hierarchical graph collaborative filtering for multimediabased recommendation,” IEEE Transactions on Computational Social Systems, vol. 11, no. 1, pp. 216-227, 2022, doi: 10.1109/tcss.2022.3226862.
M. A. Kheldouni and J. Boumhidi, “V-BERT4Rec: Enhanced sequential recommendation with multi-modal visual information,” Multimedia Tools and Applications, vol. 84, no. 11, pp. 8547-8565, 2025, doi: 10.1007/s11042-024-19277-7.
H. U. Khan, A. Naz, F. K. Alarfaj, and N. Almusallam, “A transformer-based architecture for collaborative filtering modeling in personalized recommender systems,” Scientific Reports, vol. 15, no. 1, Art. no. 24503, 2025, doi: 10.1038/s41598-025-08931-1.
Y. H. Chou, I. C. Chen, C. J. Chang, J. Ching, and Y. H. Yang, “BERT-like pre-training for symbolic piano music classification tasks,” Journal of Creative Music Systems, vol. 8, no. 1, pp. 1-19, 2024, doi: 10.5920/jcms.1064.
M. Tami, S. Masri, A. Hasasneh, and C. Tadj, “Transformer-based approach to pathology diagnosis using audio spectrogram,” Information, vol. 15, no. 5, pp. 253, 2024, doi: 10.3390/info15050253.
L. Yang, S. Wang, and B. Zhu, “Deep tensor decomposition group recommendation algorithm based on ranking,” Application Research of Computers/Jisuanji Yingyong Yanjiu, vol. 37, no. 5, pp. 1311, 2020, doi: 10.19734/j.issn.1001-3695.2019.02.006.
G. Wu, H. Qin, Q. Hu, X. Wang, and Z. Wu, “Research on large language model and its personalized recommendation,” CAAI Transactions on Intelligent Systems, vol. 19, no. 6, pp. 1351-1365, 2024, doi: 10.11992/tis.202309036.
Y. Huang, F. Zhao, X. Gui, and H. Jin, “Path-enhanced explainable recommendation with knowledge graphs,” World Wide Web, vol. 24, no. 5, pp. 1769-1789, 2021, doi: 10.1007/s11280-021-00912-4.
S. Jin, Y. Zhang, X. Li, and M. Lu, “Heterogeneous graph convolutional network for E-commerce product recommendation with adaptive denoising training,” IEEE Transactions on Consumer Electronics, vol. 70, no. 1, pp. 3259-3268, 2023, doi: 10.1109/TCE.2023.3314537.
K. Qu, K. C. Li, B. T. M. Wong, M. M. Wu, and M. Liu, “A survey of knowledge graph approaches and applications in education,” Electronics, vol. 13, no. 13, pp. 2537, 2024, doi: 10.3390/electronics13132537.
J. Chicaiza and P. Valdiviezo-Diaz, “A comprehensive survey of knowledge graph-based recommender systems: Technologies, development, and contributions,” Information, vol. 12, no. 6, pp. 232, 2021, doi: 10.3390/info12060232.
X. Hu, S. Sun, and S. Mu, “Research on intelligent recommendation of teacher training paths based on portrait technology,” e-Education Research, vol. 45, no. 2, pp. 106, 2024, doi: 10.13811/j.cnki.eer.2024.02.015.
J. Luo and Y. Zhang, “Evolution of subject knowledge graph driven by multimodal large models and its educational applications,” Modern Educational Technology, vol. 33, no. 12, pp. 76, 2023, doi: 10.3969/j.issn.1009-8097.2023.12.008.
D. Malitesta, G. Cornacchia, C. Pomo, F. A. Merra, T. Di Noia, E. Di Sciascio, et al., “Formalizing multimedia recommendation through multimodal deep learning,” ACM Transactions on Recommender Systems, vol. 3, no. 3, pp. 1-33, 2025, doi: 10.1145/3662738.
I. Aytekin, O. Dalmaz, K. Gonc, H. Ankishan, E. U. Saritas, U. Bagci, et al., “Covid-19 detection from respiratory sounds with hierarchical spectrogram transformers,” IEEE journal of biomedical and health informatics, vol. 28, no. 3, pp. 1273-1284, 2023, doi: 10.1109/JBHI.2023.3339700.
S. Masri, A. Hasasneh, M. Tami, and C. Tadj, “Exploring the impact of image-based audio representations in classification tasks using vision transformers and explainable AI techniques,” Information, vol. 15, no. 12, pp. 751, 2024, doi: 10.3390/info15120751.
S. Li and Y. Sung, “MRBERT: Pre-training of melody and rhythm for automatic music generation,” Mathematics, vol. 11, no. 4, pp. 798, 2023, doi: 10.3390/math11040798.
D. V. T. Le, L. Bigo, D. Herremans, and M. Keller, “Natural language processing methods for symbolic music generation and information retrieval: A survey,” ACM Computing Surveys, vol. 57, no. 7, pp. 1-40, 2025, doi: 10.1145/3714457.
B. Juarto and A. S. Girsang, “Neural collaborative with sentence BERT for news recommender system,” JOIV: International Journal on Informatics Visualization, vol. 5, no. 4, pp. 448-455, 2021, doi: 10.30630/joiv.5.4.678.
C. Hu, X. Sun, H. Dai, H. Zhang, and H. Liu, “Research on log anomaly detection based on sentence-BERT,” Electronics, vol. 12, no. 17, pp. 3580, 2023, doi: 10.3390/electronics12173580.
H. Abudukelimu, J. Chen, Y. Liang, A. Abulizi, and A. Yasen, “Symfornet: Application of cross-modal information correspondences based on self-supervision in symbolic music generation,” Applied Intelligence, vol. 54, no. 5, pp. 4140-4152, 2024, doi: 10.1007/s10489-024-05335-y.
N. Messina, G. Amato, A. Esuli, F. Falchi, C. Gennaro, S. Marchand-Maillet, et al., “Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 17, no. 4, pp. 1-23, 2021, doi: 10.1145/3451390.
Y. Zhong, X. Zhang, and Y. Su, “Recommendation method of teaching resources for professional music courses based on knowledge graph,” International Journal of High Speed Electronics and Systems, vol. 34, no. 04, Art. no. 2540212, 2025, doi: 10.1142/s0129156425402128.
P. Liu, Y. Cao, and L. Wang, “A multimodal fusion online music education system for universities,” Computational Intelligence and Neuroscience, vol. 2022, no. 1, Art. no. 6529110, 2022, doi: 10.1155/2022/6529110.
C. Cayari, “Popular practices for online musicking and performance: Developing creative dispositions for music education and the Internet,” Journal of Popular Music Education, vol. 5, no. 3, pp. 295-312, 2021, doi: 10.1386/jpme_00018_1.
J. Hentschel, M. Neuwirth, and M. Rohrmeier, “The annotated mozart sonatas: Score, harmony, and cadence,” Transactions of the International Society for Music Information Retrieval, vol. 4, no. 1, pp. 67-80, 2021, doi: 10.5334/TISMIR.63.
J. Hentschel, Y. Rammos, F. C. Moss, M. Neuwirth, and M. Rohrmeier, “An annotated corpus of tonal piano music from the long 19th century,” Empirical Musicology Review, vol. 18, no. 1, pp. 84-95, 2024, doi: 10.18061/emr.v18i1.8903.