Research on Intelligent Recognition Algorithm of Piano Playing Fingering and Generation of Personalized Practice Schemes
Main Article Content
Abstract
Intelligent recognition of piano fingering requires fine-grained modeling of hand motion, note events, and acoustic characteristics. Existing fingering recognition algorithms often suffer from incomplete feature extraction, weak antiinterference capability, and insufficient integration with personalized practice guidance. This study proposes a multimodal fingering recognition and practice-plan generation framework based on image, MIDI, and audio data. Highdefinition video captures hand motion trajectories, MIDI data provide note position and dynamics, and audio signals supply spectral and timbral cues. HRNet extracts 21 hand key points, Word2Vec-CBOW encodes MIDI sequences, and MFCC-based features represent audio signals. A Dynamic Channel Attention module performs multimodal feature fusion, and an improved Transformer with Position-Aware Encoding captures irregular temporal dependencies in fingering sequences. A personalized practice-generation model further combines recognition outputs, learner profiles, and teaching rules to dynamically adjust practice content, difficulty, and pace. Experiments using the MAESTRO dataset and real performance data show that the method achieves 96.2% Finger Accuracy and 94.5% Direction Accuracy, outperforming HMM, Bi-LSTM, and standard Transformer baselines. After personalized practice intervention, learners’ fingering standardization improves by 28.3% and practice efficiency by 32.1%. The framework supports multimodal signal processing, temporal sequence modeling, and intelligent music-training systems.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
L. Sun, “Historical Traces in the RCM Piano Score Archive: Evolution of Pedagogical Approaches Through Editions,” Cultural and Religious Studies, vol. 13, no. 4, pp. 209-217, 2025, doi: 10.17265/2328-2177/2025.04.008.
Y. Zhang, “Status and Reform of Piano Teaching in Public Art Education in Colleges and Universities under the Background of Deep Learning,” Archives des Sciences, vol. 74, no. s2, pp. 36-42, 2024, doi: 10.62227/as/74s25.
J. Woosung, L. Eunjoo, and C. Suah, “A technique to support the personalized learning based on the log data of piano chords practicing,” The Journal of the Institute of Internet, Broadcasting and Communication, vol. 23, no. 1, pp. 191-201, 2023.
H. Wang and L. Zhu, “The Application of XAPT-based Blended Learning in Piano Courses for Non-Piano Majors in Qinghai Region,” International Journal of Sociologies and Anthropologies Science Reviews, vol. 5, no. 1, pp. 395-406, 2025, doi: 10.60027/ijsasr.2025.5370.
M. Yin and T. Sondhiratna, “The Perspectives On Piano Teaching Strategies At Qingdao Art School in Shandong Province,” International Journal of Sociologies and Anthropologies Science Reviews, vol. 4, no. 2, pp. 527-534, 2024, doi: 10.60027/ijsasr.2024.4481.
Q. Qiu, “Sentiment analysis based on multimodal feature fusion,” IET Conference Proceedings, vol. 2024, no. 24, pp. 497-503, 2025, doi: 10.1049/icp.2024.4566.
K. Zhou and H. Yan, “MMFFNet: Multimodal Feature Fusion Network for RGB-T Crowd Counting,” IEEE Transactions on Instrumentation and Measurement, vol. 74, pp. 1-12, 2025, doi: 10.1109/tim.2025.3618732.
L. X. Qin, H. M. Sun, X. M. Duan, et al., “MFCNet: Multimodal Feature Fusion Network for RGB-T Vehicle Density Estimation,” IEEE Internet of Things Journal, vol. 12, no. 4, pp. 4207-4219, 2025, doi: 10.1109/JIOT.2024.3483175.
W. Yan, W. Liu, Q. Zhang, et al., “Multisource Multimodal Feature Fusion for Small Leak Detection in Gas Pipelines,” IEEE Sensors Journal, vol. 24, no. 2, pp. 1857-1865, 2024, doi: 10.1109/JSEN.2023.3337228.
Y. Lin, Q. Chen, F. Wang, et al., “Automatic Modulation Classification Based on Efficient Multimodal Feature Fusion,” Mobile Networks and Applications, vol. 30, no. 3-4, pp. 628-637, 2025, doi: 10.1007/s11036-025-02487-0.
J. Hao, “Educational resource allocation optimisation driven by multimodal feature fusion,” International Journal of Information and Communication Technology, vol. 26, no. 31, pp. 105-125, 2025, doi: 10.1504/IJICT.2025.148129.
E. Ka, S. Go, M. Kwak, et al., “Solar Power Generation Forecasting via Multimodal Feature Fusion (Student ABSTRACT ),” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 21, pp. 23530-23532, 2024, doi: 10.1609/aaai.v38i21.30459.
H. Han, Y. Meng, X. Wu, et al., “A Transfer Learning-Based Multimodal Feature Fusion Model for Bearing Fault Diagnosis,” IEEE Transactions on Instrumentation and Measurement, vol. 74, pp. 1-13, 2025, doi: 10.1109/TIM.2025.3558745.