Design of University Instrumental Music Performance Movement Recognition and Accurate Feedback System Integrating OpenPose and LSTM
Main Article Content
Abstract
Accurate recognition of fine-grained human movements with low-latency feedback is fundamental to intelligent perception systems and real-time human–machine interaction in modern engineering applications. This study presents an end-to-end motion recognition and feedback framework for university instrumental music performance by integrating OpenPose-based pose estimation with a three-layer bidirectional long short-term memory (BiLSTM) network. The proposed approach extracts normalized skeletal keypoint sequences from performance videos and exploits bidirectional temporal modeling to capture both local motion evolution and long-range spatiotemporal dependencies. An adaptive feedback module further quantifies posture deviations through coordinate distance analysis and automatically generates structured correction instructions based on predefined performance standards. Experimental results demonstrate that the three-layer BiLSTM architecture achieves a motion recognition accuracy of 96.3%, while the optimized system reduces end-to-end latency to 48.3 ms and maintains 98.2% overall accuracy under standard teaching conditions. The framework also exhibits robust performance across challenging environments, including low illumination and complex backgrounds. By combining vision-based sensing with temporal sequence learning, the proposed method provides an effective solution for objective motion evaluation and intelligent feedback generation, while offering valuable methodological references for real-time signal interpretation, human-centered electromagnetic sensing, and next-generation intelligent monitoring systems requiring precise dynamic feature extraction.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
E. Sanchez-Escribano, F. Gertrudix, and A. Bautista, “Analyzing instrumental music education models: A fourdimension tool,” Arts Education Policy Review, vol. 125, no. 3, pp. 187-197, 2024, doi: 10.1080/10632913.2022.2041139.
E. Concina, “Effective music teachers and effective music teaching today: A systematic review,” Education Sciences, vol. 13, no. 2, pp. 107-128, 2023, doi: 10.3390/educsci13020107.
G. G. Artıktay, “Cognitive neuroscience and music education: Relationships and interactions,” International Journal of Educational Spectrum, vol. 6, no. 1, pp. 91-119, 2024, doi: 10.47806/ijesacademic.1402953.
T. Ma, I. C. Sanchis, G. R. Santana, and Y. Jiang, “Neuroplasticity Mechanisms in Early Childhood Piano Education: A Literature Review from the Perspective of Educational Neuroscience,” Journal of Sociology and Education, vol. 1, no. 1, pp. 167-175, 2025, doi: 10.63887/jse.2025.1.1.19.
G. Y. Duztaban, “THE PIANO EDUCATION IN EARLY CHILDHOOD: A COMPARATIVE ANALYSIS OF MUSIC EDUCATION MODELS,” LOKUM Sanat ve Tasarım Dergisi, vol. 3, no. 2, pp. 230-252, 2025, [Online]. Available: https://izlik.org/JA46DY98FX.
F. Roggio, B. Trovato, M. Sortino, and G. Musumeci, “A comprehensive analysis of the machine learning pose estimation models used in human movement and posture analyses: A narrative review,” Heliyon, vol. 10, no. 21, Art. no. e39977-e39989, 2024, doi: 10.1016/j.heliyon.2024.e39977.
Y. Liu, T. Zhang, Z. Li, and L. Deng, “Deep learning-based standardized evaluation and human pose estimation: A novel approach to motion perception,” Traitement du Signal, vol. 40, no. 5, pp. 2313-2320, 2023, doi: 10.18280/ts.400549.
S. Kulkarni, S. Deshmukh, F. Fernandes, A. Patil, and V. Jabade, “Poseanalyser: A survey on human pose estimation,” SN Computer Science, vol. 4, no. 2, pp. 136, 2023, doi: 10.1007/s42979-022-01567-2.
K. Esaki and K. Nagao, “An efficient immersive self-training system for hip-hop dance performance with automatic evaluation features,” Applied Sciences, vol. 14, no. 14, pp. 5981, 2024, doi: 10.3390/app14145981.
X. Ma, “Deep learning models combining stereo vision for dance movement evaluation,” International Journal of Information and Communication Technology, vol. 26, no. 11, pp. 69-85, 2025, doi: 10.1504/ijict.2025.146102.
Y. Zhang, “Virtual Concerts in Learning Oboe-Played Chinese Folk Music: Impact on Performance Proficiency, Perceived Aesthetic Qualities, and Students ’ Motivation,” International Review of Research in Open and Distributed Learning, vol. 26, no. 3, pp. 61-82, 2025, doi: 10.19173/irrodl.v26i3.8358.
N. Moura and S. Serra, “Saxophone players ’ self-perceptions about body movement in music performing and learning: An interview study,” Music Perception: An Interdisciplinary Journal, vol. 41, no. 3, pp. 199-216, 2024, doi: 10.1525/mp.2024.41.3.199.
L. Zhou, “Construction and Practice of Erhu Performance Course for Musicology Majors in Local Universities—A Case Study of Sichuan Minzu College,” Frontiers in Educational Research, vol. 7, no. 2, pp. 71-75, 2024, doi: 10.25236/FER.2024.070211.
N. Hassan, A. S. M. Miah, and J. Shin, “A deep bidirectional LSTM model enhanced by transfer-learning-based feature extraction for dynamic human activity recognition,” Applied Sciences, vol. 14, no. 2, pp. 603, 2024, doi: 10.3390/app14020603.
N. Zerrouki, F. Harrou, A. Houacine, R. Bouarroudj, M. Y. Cherifi, A. D. A. Zouina, et al., “Deep learning for hand gesture recognition in virtual museum using wearable vision sensors,” IEEE Sensors Journal, vol. 24, no. 6, pp. 8857-8869, 2024, doi: 10.1109/JSEN.2024.3354784.
J. Wu, P. Ren, B. Song, R. Zhang, C. Zhao, and X. Zhang, “Data glove-based gesture recognition using CNN-BiLSTM model with attention mechanism,” PLoS One, vol. 18, no. 11, Art. no. e0294174-e0294195, 2023, doi: 10.1371/journal.pone.0294174.
S. Sahoo, “Sensor Fusion and Virtual Sensor Design for Enhanced Multi-Sensor Data Accuracy in Autonomous Systems,” International Journal on Smart & Sustainable Intelligent Computing, vol. 1, no. 2, pp. 21-39, 2024, doi: 10.63503/j.ijssic.2024.31.
Y. Zhang, B. Zhang, C. Shen, H. Liu, J. Huang, K. Tian, et al., “Review of the field environmental sensing methods based on multi-sensor information fusion technology,” International Journal of Agricultural and Biological Engineering, vol. 17, no. 2, pp. 1-13, 2024, doi: 10.25165/j.ijabe.20241702.8596.
P. Lalwani and R. Ganeshan, “A novel CNN-BiLSTM-GRU hybrid deep learning model for human activity recognition,” International Journal of Computational Intelligence Systems, vol. 17, no. 1, pp. 278-297, 2024, doi: 10.1007/s44196-024-00689-0.
Z. Chen, Q. Xie, and W. Jiang, “Hybrid Deep Learning Models for Tennis Action Recognition: Enhancing Professional Training Through CNN-BiLSTM Integration,” Concurrency and Computation: Practice and Experience, vol. 37, no. 6-8, Art. no. e70029, 2025, doi: 10.1002/cpe.70029.
A. Tharatipyakul, T. Srikaewsiew, and S. Pongnumkul, “Deep learning-based human body pose estimation in providing feedback for physical movement: A review,” Heliyon, vol. 10, no. 17, Art. no. e36589-e36604, 2024, doi: 10.1016/j.heliyon.2024.e36589.
W. Zhang, “Dynamic pose recognition based on deep learning: Developing a CNN model for choral conductor pose recognition,” Journal of Computational Methods in Sciences and Engineering, vol. 25, no. 4, pp. 3856-3874, 2025, doi: 10.1177/14727978251323068.
M. Sui, L. Jiang, T. Lyu, H. Wang, L. Zhou, and Chen P.et al, “Application of deep learning models based on efficientdet and openpose in user-oriented motion rehabilitation robot control,” Journal of Intelligence Technology and Innovation, vol. 2, no. 3, pp. 47-77, 2024, doi: 10.30212/JITI.202402.012.
L. Xiao, Y. Cao, Y. Gai, E. Khezri, J. Liu, and M. Yang, “Recognizing sports activities from video frames using deformable convolution and adaptive multiscale features,” Journal of Cloud Computing, vol. 12, no. 1, pp. 167-186, 2023, doi: 10.1186/s13677-023-00552-1.
S. Lin and W. Hou, “Efficient Sampling of Two-Stage Multi-Person Pose Estimation and Tracking from Spatiotemporal,” Applied Sciences, vol. 14, no. 6, Art. no. 2238, 2024, doi: 10.3390/app14062238.
L. Bellier, A. Llorens, D. Marciano, A. Gunduz, G. Schalk, P. Brunner, et al., “Music can be reconstructed from human auditory cortex activity using nonlinear decoding models,” PLoS biology, vol. 21, no. 8, Art. no. e3002176-e3002186, 2023, doi: 10.1371/journal.pbio.3002176.
L. Wang, “Neural Networks and Ensemble Model to Automatic Music Coordination: A Performance Comparison,” Information Technology and Control, vol. 54, no. 2, pp. 520-535, 2025, doi: 10.5755/j01.itc.54.2.36737.
N. Mohammadi Foumani, L. Miller, C. W. Tan, G. I. Webb, G. Forestier, and M. Salehi, “Deep learning for time series classification and extrinsic regression: A current survey,” ACM Computing Surveys, vol. 56, no. 9, pp. 1-45, 2024, doi: 10.1145/3649448.
X. Li, Z. Luo, J. Lv, C. Yang, S. Yan, J. Xu, et al., “m5U-HybridNet: Integrating an RNA Foundation Model with CNN Features for Accurate Prediction of 5-Methyluridine Modification Sites,” Journal of Chemical Information and Modeling, vol. 65, no. 15, pp. 8079-8096, 2025, doi: 10.1021/acs.jcim.5c01237.
X. Zhao, L. Wang, Y. Zhang, X. Han, M. Deveci, and M. Parmar, “A review of convolutional neural networks in computer vision,” Artificial Intelligence Review, vol. 57, no. 4, pp. 99-141, 2024, doi: 10.1007/s10462-024-10721-6.
Q. Lei, H. Li, H. Zhang, J. Du, and S. Gao, “Multi-skeleton structures graph convolutional network for action quality assessment in long videos,” Applied Intelligence, vol. 53, no. 19, pp. 21692-21705, 2023, doi: 10.1007/s10489-023-04613-5.
Y. Han, L. Han, C. Zeng, and W. Zhao, “The innovation path of VR technology integration into music classroom teaching in colleges and universities,” Scientific Reports, vol. 15, no. 1, pp. 12200-12216, 2025, doi: 10.1038/s41598-025-97003-5.
Y. Jiang and D. Dumlavwalla, “Pre-college piano teaching practice in China and the United States: An international comparative study from the piano teachers ’ perspectives,” International Journal of Music Education, vol. 43, no. 3, pp. 519-533, 2025, doi: 10.1177/02557614231208236.
Y. Li, “Interactive methods for improving musical literacy among students with preschool education Majors at Teacher Training Universities: The effectiveness of the Kodály method,” Education and Information Technologies, vol. 28, no. 10, pp. 12807-12821, 2023, doi: 10.1007/s10639-023-11636-5.
L. R. de Bruin, “Instrumental music teachers ’ development of feedback across the lifespan: A qualitative study,” International Journal of Music Education, vol. 42, no. 1, pp. 32-46, 2024, doi: 10.1177/02557614231151445.
L. R. de Bruin, “Feedback in the instrumental music lesson: A qualitative study,” Psychology of Music, vol. 51, no. 4, pp. 1259-1274, 2023, doi: 10.1177/03057356221135668.