Real-Time Assessment of Piano Sight-Reading Ability in College Students Using CRNN
Main Article Content
Abstract
This study proposes a real-time piano sight-reading assessment framework based on a multi-task convolutional recurrent neural network (MT-CRNN) for objective and intelligent evaluation of college students’ performance ability. The proposed architecture integrates a shared convolutional encoder with parallel decoding branches to simultaneously perform note transcription and musical expression parameter regression. A music rule-based feature fusion module is further designed to align discrete note events with continuous acoustic features, generating a unified representation for multidimensional performance analysis. Based on this representation, an evaluation system is constructed to assess pitch accuracy, rhythmic stability, dynamic control, phrase coherence, and overall musicality. Experimental validation was conducted on a dedicated piano sight-reading dataset containing 400 performance recordings from undergraduate music students. Results demonstrate that the proposed method achieves pitch and duration F1-scores of 95.3% and 94.1 %, respectively, outperforming conventional transcription-based approaches. The generated evaluation scores exhibit strong agreement with expert assessments, with Spearman correlation coefficients reaching 0.84 for dynamic control, 0. 82 for phrase coherence, and 0.79 for overall musicality. Furthermore, the system processes an average 46.8-second performance in only 2.75 seconds, satisfying real-time feedback requirements. The proposed framework provides an effective engineering solution for automated music performance assessment by combining deep learning, acoustic signal analysis, and knowledge-guided evaluation mechanisms, offering practical support for intelligent music education systems.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
R. Rui, M. S. Amran, and N. M. Nasri, “Piano sight-reading teaching and music education: A systematic literature review,” JOURNAL OF INFRASTRUCTURE, POLICY AND DEVELOPMENT, vol. 9, no. 1, Art. no. 10391, 2025, doi: 10.24294/jipd10391.
N. Jia and P. Roongruang, “Training Students’ Sight-Reading Skill for Piano Teaching in Colleges Under Multicultural Background in China,” Journal of Modern Learning Development, vol. 8, no. 3, pp. 409-416, 2023, [Online]. Available: https://so06.tci-thaijo.org/index.php/jomld/article/view/259443.
V. Phanichraksaphong and W. H. Tsai, “Automatic assessment of piano performances using timbre and pitch features,” Electronics, vol. 12, no. 8, pp. 1791, 2023, doi: 10.3390/electronics12081791.
X. Zhao, Y. Wang, and X. Cai, “A ResNet-based audio-visual fusion model for piano skill evaluation,” Applied Sciences, vol. 13, no. 13, pp. 7431, 2023, doi: 10.3390/app13137431.
J. Park, J. Kim, J. M. Park, A. Choi, W. S. Li, J. Park, et al., “Piano performance evaluation dataset with multilevel perceptual features,” Scientific Reports, vol. 14, no. 1, Art. no. 23002, 2024, doi: 10.1038/s41598-024-73810-0.
D. Mao and S. Liu, “Quantitative Evaluation of Piano Performance Technique and Style Based on Continuous Discretization Method,” Applied Mathematics and Nonlinear Sciences, vol. 9, no. 1, pp. 1, 2024, doi: 10.2478/amns.2023.2.00093.
V. Phanichraksaphong and W. H. Tsai, “Automatic evaluation of piano performances for STEAM education,” Applied Sciences, vol. 11, no. 24, Art. no. 11783, 2021, doi: 10.3390/app112411783.
W. Wang, J. Pan, H. Yi, Z. Song, and M. Li, “Audio-based piano performance evaluation for beginners with convolutional neural network and attention mechanism,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, no. 1, pp. 1119-1133, 2021, doi: 10.1109/taslp.2021.3061267.
Z. Xiao, X. Chen, and L. Zhou, “Polyphonic piano transcription based on graph convolutional network,” Signal Processing, vol. 212, no. 1, Art. no. 109134, 2023, doi: 10.1016/j.sigpro.2023.109134.
M. Li, “Design and implementation of piano audio automatic music transcription algorithm based on convolutional neural network,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2025, no. 1, pp. 26, 2025, doi: 10.1186/s13636-025-00412-7.
J. Dai, Q. Zheng, Y. Wang, Q. Shan, J. Wan, and W. Zhang, “Multi-Feature Fusion for Automatic Piano Transcription Based on Mel Cyclic and STFT Spectrograms,” Electronics, vol. 14, no. 23, pp. 4720, 2025, doi: 10.3390/electronics14234720.
R. Guo and Y. Zhu, “Research on the Recognition of Piano-Playing Notes by a Music Transcription Algorithm,” Journal of Advanced Computational Intelligence and Intelligent Informatics, vol. 29, no. 1, pp. 152-157, 2025, doi: 10.20965/jaciii.2025.p0152.
M. Zhang, “Advancing deep learning for expressive music composition and performance modeling,” Scientific Reports, vol. 15, no. 1, Art. no. 28007, 2025, doi: 10.1038/s41598-025-13064-6.
H. Zhang, S. Chowdhury, C. E. Cancino-Chacón, J. Liang, S. Dixon, and G. Widmer, “Dexter: Learning and controlling performance expression with diffusion models,” Applied Sciences, vol. 14, no. 15, pp. 6543, 2024, doi: 10.3390/app14156543.
V. Sareen and K. R. Seeja, “Speech Emotion Recognition using Mel Spectrogram and Convolutional Neural Networks (CNN),” Procedia Computer Science, vol. 258, no. 1, pp. 3693-3702, 2025, doi: 10.1016/j.procs.2025.04.624.
P. Rawat, M. Bajaj, S. Vats, and V. Sharma, “A comprehensive study based on MFCC and spectrogram for audio classification,” Journal of Information and Optimization Sciences, vol. 44, no. 6, pp. 1057-1074, 2023, doi: 10.47974/JIOS-1431.
J. H. Wang, P. T. Le, S. J. Kuo, T. C. Tai, K. C. Li, S. L. Chen, et al., “Audio pre-processing and beam-forming implementation on embedded systems,” Electronics, vol. 13, no. 14, pp. 2784, 2024, doi: 10.3390/electronics13142784.
Barkovcka OYu and A. O. Gavrashenko, “Research of the impact of noise reduction methods on the quality of audio signal recovery,” Informatsiyno-Keryyuchi Cictemi na Zaliznichnomy Trancporti, vol. 29, no. 3, pp. 57-65, 2024, doi: 10.18664/ikszt.v29i3.313606.
T. Kawamura, Y. Kinoshita, N. Ono, and R. Scheibler, “Acoustic scene classification using inter-and intra-subarray spatial features in distributed microphone array,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, pp. 65, 2024, doi: 10.1186/s13636-024-00386-y.
C. Asuai, A. P. Arinomor, C. T. Atumah, I. F. Kowhoro, and D. E. Ogheneochuko, “Hybrid CNN-LSTM Architectures for Deepfake Audio Detection Using Mel Frequency Cepstral Coefficients and Spectogram Analysis,” American Journal of Mathematical and Computer Modelling, vol. 10, no. 3, pp. 98-109, 2025, doi: 10.11648/j.ajmcm.20251003.12.
M. Swain, B. Maji, P. Kabisatpathy, and A. Routray, “A DCRNN-based ensemble classifier for speech emotion recognition in Odia language,” Complex & Intelligent Systems, vol. 8, no. 5, pp. 4237-4249, 2022, doi: 10.1007/s40747-022-00713-w.
S. Madanian, O. Adeleye, J. M. Templeton, T. Chen, C. Poellabauer, E. Zhang, et al., “A multi-dilated convolution network for speech emotion recognition,” Scientific Reports, vol. 15, no. 1, pp. 8254, 2025, doi: 10.1038/s41598-025-92640-2.
Y. Feng, “Intelligent speech recognition algorithm in multimedia visual interaction via BiLSTM and attention mechanism,” Neural Computing and Applications, vol. 36, no. 5, pp. 2371-2383, 2024, doi: 10.1007/s00521-023-08959-2.
Y. L. Chen, N. C. Wang, J. F. Ciou, and R. Q. Lin, “Combined bidirectional long short-term memory with mel-frequency cepstral coefficients using autoencoder for speaker recognition,” Applied Sciences, vol. 13, no. 12, pp. 7008, 2023, doi: 10.3390/app13127008.
F. Simonetta, F. Avanzini, and S. Ntalampiras, “A perceptual measure for evaluating the resynthesis of automatic music transcriptions,” Multimedia Tools and Applications, vol. 81, no. 22, pp. 32371-32391, 2022, doi: 10.1007/s11042-022-12476-0.
P. Wang and N. Dai, “Processing piano audio: Research on an automatic transcription model for sound signals,” Journal of Measurements in Engineering, vol. 13, no. 1, pp. 130-139, 2025, doi: 10.21595/jme.2024.24345.
K. Shibata, E. Nakamura, and K. Yoshii, “Non-local musical statistics as guides for audio-to-score piano transcription,” Information Sciences, vol. 566, no. 1, pp. 262-280, 2021, doi: 10.1016/j.ins.2021.03.014.
B. Bhattarai and J. Lee, “A comprehensive review on music transcription,” Applied Sciences, vol. 13, no. 21, Art. no. 11882, 2023, doi: 10.3390/app132111882.
G. Jones and A. Friberg, “Probing the underlying principles of dynamics in piano performances using a modelling approach,” Frontiers in Psychology, vol. 14, no. 1, Art. no. 1269715, 2023, doi: 10.3389/fpsyg.2023.1269715.
M. Nusseck, I. Czedik-Eysenberg, C. Spahn, and C. Reuter, “Associations between ancillary body movements and acoustic parameters of pitch, dynamics and timbre in clarinet playing,” Frontiers in Psychology, vol. 13, no. 1, Art. no. 885970, 2022, doi: 10.3389/fpsyg.2022.885970.