Research on Deep Learning-Based Methods for Emotional Expression Recognition and Style Analysis in Piano Performance
Main Article Content
Abstract
Emotional expression recognition in piano performance requires fine-grained modeling of audio timbre, dynamic touch, rhythm fluctuation, and performance style. Existing methods based on single audio features or shallow statistics often fail to capture subtle emotional transitions and stylistic differences. This study proposes a deep-learning-based framework integrating multimodal feature fusion, temporal modeling, attention mechanisms, and emotion-style collaborative learning. Piano audio signals are converted into time-frequency representations through short-time Fourier transform and Mel-scale mapping, while MIDI control information encodes velocity, duration, and beat-offset features. Audio spectrogram features and MIDI control vectors are fused to construct a unified representation. A convolutional network extracts local time-frequency patterns related to accents and touch variations, and a temporal modeling module captures emotional evolution across performance sequences. An energy- and velocity-guided attention mechanism enhances key musical segments, including climaxes, strong beats, and ornaments. In addition, a style representation branch models rhythmic stability, dynamic fluctuation, and tempo evolution, while cross-task consistency constraints couple emotion recognition and style analysis. Experiments on paired audio–MIDI performance data show that the proposed method achieves 0.91 accuracy and 0.90 F1 score, improving by approximately 9% and 10% over unimodal baselines. The framework provides an audio signal processing and multimodal sequence-analysis method for intelligent music-performance evaluation.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
Y. Zhang, Y. Feng, and Q. Ding, “A Multi-kernel Convolutional Neural Network for Multi-emotion Classification of Music Lyrics Based on Word2Vec Word Embedding,” Science Technology and Engineering, vol. 24, no. 20, pp. 8598-8605, 2024.
C. Xiao, “A Brief Analysis of the Path to Improve the Literacy and Ability of Piano Art Instructors,” Modern Education Frontiers, vol. 6, no. 9, pp. 15-17, 2025.
Q. Li, Z. Wu, and X. Guan, “A method for generating piano fingerings by combining deep musical score features,” Journal of Intelligent Systems, vol. 18, no. 6, pp. 1287-1294, 2023.
Y. Li, L. Wang, Y. Xue, et al., “Analysis and research on intelligent music generation system based on short-time Fourier transform,” Journal of Intelligent Systems, vol. 20, no. 3, pp. 750-760, 2025.
Y. Xu, M. Cheng, H. Shao, et al., “The relationship between music learning and phonological awareness: evidence from metaanalysis,” Psychological Technology and Application, vol. 11, no. 12, pp. 705-721, 2023.
J. Wang and R. Huang, “Music Emotion Recognition Based on Wide-Depth Learning Network,” Journal of East China University of Science and Technology (Natural Science Edition), vol. 48, no. 3, pp. 373-380, 2022.
Y. Zhu, T. Feng, M. Zhang, et al., “Convolutional speech emotion recognition network based on incremental method,” Journal of Shanghai University (Natural Science Edition), vol. 29, no. 1, pp. 24-40, 2023.