Vocal Learner Ability Profiles Integrating Graph Neural Networks
Main Article Content
Abstract
Existing competency assessments often evaluate individual skills independently and fail to represent the structural dependencies between professional knowledge and practical performance. This paper applies Graph Neural Networks to construct dynamic ability profiles for vocal learners based on knowledge-structure modeling. Practice records and curriculum data are integrated into a heterogeneous graph containing learners, skills, and knowledge points. An R-GCN first generates initial embeddings through relation-specific message passing, preserving the semantic effects of different edge types. A Temporal GAT then applies temporal encoding to dynamic graph sequences to capture the evolution of learner abilities over time. The final time-series ability embedding is mapped to five dimensions, including pitch, rhythm, breath, resonance, and emotional expression. Experimental results show that structural modeling achieves average node reconstruction accuracy of at least 0.81 and relationship prediction F1 of at least 0.80. A dynamic embedding similarity slope of 0.82 and ability-change direction cosine of 0.85 indicate high temporal consistency in tracking ability evolution. The proposed framework supports structured inference of learner competence and provides a technical basis for personalized vocal-training intervention.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
J. N. Audet, M. Couture, and E. D. Jarvis, “Songbird species that display more-complex vocal learning are better problem-solvers and have larger brains,” Science, vol. 381, no. 6663, pp. 1170-1175, 2023, doi: 10.1126/science.adh3428.
A. Rybner, E. T. Jessen, M. D. Mortensen, S. N. Larsen, R. Grossman, N. Bilenberg, et al., “Vocal markers of autism: Assessing the generalizability of machine learning models,” Autism Research, vol. 15, no. 6, pp. 1018-1030, 2022, doi: 10.1002/aur.2721.
Z. You, B. Han, Z. Shi, M. Zhao, S. Du, J. Yan, et al., “Vocal cord leukoplakia classification using deep learning models in white light and narrow band imaging endoscopy images,” Head & Neck, vol. 45, no. 12, pp. 3129-3145, 2023, doi: 10.1002/hed.27543.
T. O’Rourke, P. T. Martins, R. Asano, R. O. Tachibana, K. Okanoya, and C. Boeckx, “Capturing the effects of domestication on vocal learning complexity,” Trends in Cognitive Sciences, vol. 25, no. 6, pp. 462-474, 2021, doi: 10.1016/j.tics.2021.05.002.
Q. Zhao, Y. He, Y. Wu, D. Huang, Y. Wang, C. Sun, et al., “Vocal cord lesions classification based on deep convolutional neural network and transfer learning,” Medical Physics, vol. 49, no. 1, pp. 432-442, 2022, doi: 10.1002/mp.15371.
J. A. Cahill, J. Armstrong, A. Deran, C. J. Khoury, B. Paten, D. Haussler, et al., “Positive selection in noncoding genomic regions of vocal learning birds is associated with genes implicated in vocal learning and speech functions in humans,” Genome Research, vol. 31, no. 11, pp. 2035-2049, 2021, doi: 10.1101/gr.275989.121.
C. W. Espinola, J. C. Gomes, J. M. S. Pereira, and W. P. dos Santos, “Vocal acoustic analysis and machine learning for the identification of schizophrenia,” Research on Biomedical Engineering, vol. 37, no. 1, pp. 33-46, 2021, doi: 10.1007/s42600-020-00097-1.
Y. Shi, “The use of mobile internet platforms and applications in vocal training: Synergy of technological and pedagogical solutions,” Interactive Learning Environments, vol. 31, no. 6, pp. 3780-3791, 2023, doi: 10.1080/10494820.2021.1943456.
É. R. Santana, L. Lopes, and R. M. de Moraes, “Recognition of the effect of vocal exercises by fuzzy triangular naive Bayes, a machine learning classifier: A preliminary analysis,” Journal of Voice, vol. 39, no. 2, pp. 560.e21-560.e30, 2025, doi: 10.1016/j.jvoice.2022.10.001.
Y. Wu, E. D. Jarvis, and A. Sarkar, “Bayesian semiparametric Markov renewal mixed models for vocalization syntax,” Biostatistics, vol. 25, no. 3, pp. 648-665, 2024, doi: 10.1093/biostatistics/kxac050.
Z. Ma, P. Ruannakarn, and A. Jansaeng, “The role of blended instructional models based on deep learning theory in enhancing students’ autonomous learning ability in Chinese Higher Vocational Colleges,” Journal of Statistics Applications and Probability, vol. 13, no. 3, pp. 947-959, 2024, doi: 10.18576/jsap/130309.
M. D. Pawar and R. D. Kokate, “Convolution neural network based automatic speech emotion recognition using Mel-frequency Cepstrum coefficients,” Multimedia Tools and Applications, vol. 80, no. 10, pp. 15563-15587, 2021, doi: 10.1007/s11042-020-10329-2.
S. Mustafa, A. Khan, S. Hussain, M. Z. Jhandir, R. Kazmi, and I. S. Bajwa, “Automatic speech emotion recognition using Mel frequency cepstrum co-efficient and machine learning technique,” Pakistan Journal of Engineering and Technology, vol. 4, no. 1, pp. 124-130, 2021, doi: 10.51846/vol4iss1pp124-130.
Y. Fang, X. Li, R. Ye, X. Tan, P. Zhao, and M. Wang, “Relation-aware graph convolutional networks for multirelational network alignment,” ACM Transactions on Intelligent Systems and Technology, vol. 14, no. 2, pp. 1-23, 2023, doi: 10.1145/3579827.
M. Zhang, H. Liu, C. Chen, G. Gao, H. Li, and J. Zhao, “AccessFixer: Enhancing GUI accessibility for low vision users with R-GCN model,” IEEE Transactions on Software Engineering, vol. 50, no. 2, pp. 173-189, 2023, doi: 10.1109/TSE.2023.3337421.