A Conv-TCN Model for Extracting Dynamic Rhythm Paths in Expressive Piano Performances
Main Article Content
Abstract
Accurate characterization of expressive rhythm evolution remains a challenging task because conventional methods primarily rely on static descriptors and fail to capture long-range temporal dependencies. This study proposes a Conv-TCN framework for extracting dynamic rhythm paths from piano performance data by integrating one-dimensional convolutional neural networks with temporal convolutional networks. Temporal features are first extracted from MIDI sequences and encoded to represent local rhythmic structures, while dilated temporal convolutions model long-term dependencies and gradual expressive variations. A multi-scale fusion strategy further combines rhythmic information across different temporal resolutions to generate continuous and interpretable rhythm trajectories. Experimental results demonstrate that the proposed method achieves an MSE of 0.032, smoothness of 0.020, KL divergence of 0.052, and similarity of 0.83 under noisy conditions, outperforming several representative sequence modeling approaches. Beyond music performance analysis, the proposed temporal feature extraction and dynamic trajectory modeling strategy provides a methodological reference for processing complex non-stationary signals, with potential value for electromagnetic signal characterization, antenna measurement data analysis, and propagation-related time-series interpretation.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
W. G. Chen, J. R. Iversen, M. H. Kao, P. Loui, A. D. Patel, R. J. Zatorre, et al., “Music and brain circuitry: Strategies for strengthening evidence-based research for music-based interventions,” Journal of Neuroscience, vol. 42, no. 45, pp. 8498-8507, 2022, doi: 10.1523/JNEUROSCI.1135-22.2022.
S. Ji, X. Yang, and J. Luo, “A survey on deep learning for symbolic music generation: Representations, algorithms, evaluations, and challenges,” ACM Computing Surveys, vol. 56, no. 1, pp. 1-39, 2023, doi: 10.1145/3597493.
C. Weiß, F. Zalkow, V. Arifi-Müller, M. Müller, H. V. Koops, A. Volk, et al., “Schubert Winterreise dataset: A multimodal scenario for music analysis,” Journal on Computing and Cultural Heritage (JOCCH), vol. 14, no. 2, pp. 1-18, 2021, doi: 10.1145/3429743.
R. Su, “Piano teaching-assisted beat recognition based on spatio-temporal two-branch attention,” International Journal of Information and Communication Technology, vol. 26, no. 5, pp. 100-116, 2025, doi: 10.1504/IJICT.2025.145149.
Y. Su and Y. Wang, “Optimization of music education strategy guided by the temporal-difference reinforcement learning algorithm,” Soft Computing, vol. 28, no. 13, pp. 8279-8291, 2024, doi: 10.1007/s00500-024-09631-0.
Z. Liu, “AI-Driven classification and trend analysis of piano music genres using large language models,” International Journal of Information and Communication Technology, vol. 26, no. 12, pp. 32-48, 2025, doi: 10.1504/IJICT.2025.146164.
Y. Jiang and Z. Sun, “Multimodal Sentiment Perception for Intelligent Music Generation Using Improved Transformer Architectures,” Informatica, vol. 49, no. 5, pp. 147-166, 2025, doi: 10.31449/inf.v49i5.6864.
L. Wang, Z. Zhao, H. Liu, J. Pang, Y. Qin, and Q. Wu, “A review of intelligent music generation systems,” Neural Computing and Applications, vol. 36, no. 12, pp. 6381-6401, 2024, doi: 10.1007/s00521-024-09418-2.
Z. Xiao, X. Chen, and L. Zhou, “Music performance style transfer for learning expressive musical performance,” Signal, Image and Video Processing, vol. 18, no. 1, pp. 889-898, 2024, doi: 10.1007/s11760-023-02788-5.
D. V. T. Le, L. Bigo, D. Herremans, and M. Keller, “Natural language processing methods for symbolic music generation and information retrieval: A survey,” ACM Computing Surveys, vol. 57, no. 7, pp. 1-40, 2025, doi: 10.1145/3714457.
J. Min, Z. Gao, L. Wang, and A. Zhang, “Application research of short-time Fourier transform in music generation based on the parallel WaveGan system,” IEEE Transactions on Industrial Informatics, vol. 20, no. 9, pp. 10770-10778, 2024, doi: 10.1109/TII.2024.3397344.
P. Georges and A. Seckin, “Music information visualization and classical composers discovery: An application of network graphs, multidimensional scaling, and support vector machines,” Scientometrics, vol. 127, no. 5, pp. 2277-2311, 2022, doi: 10.1007/s11192-022-04331-8.
F. Liu, D. L. Chen, R. Z. Zhou, S. Yang, and F. Xu, “Self-Supervised music motion synchronization learning for music-driven conducting motion generation,” Journal of Computer Science and Technology, vol. 37, no. 3, pp. 539-558, 2022, doi: 10.1007/s11390-022-2030-z.
H. Jia, “Piano performance techniques and musical expressiveness,” Pacific International Journal, vol. 6, no. 4, pp. 34-37, 2023, doi: 10.55014/pij.v6i4.461.
L. Peng, “Piano players’ intonation and training using deep learning and mobilenet architecture,” Mobile Networks and Applications, vol. 28, no. 6, pp. 2182-2190, 2023, doi: 10.1007/s11036-023-02175-x.
Y. Ghatas, M. Fayek, and M. Hadhoud, “A hybrid deep learning approach for musical difficulty estimation of piano symbolic music,” Alexandria Engineering Journal, vol. 61, no. 12, pp. 10183-10196, 2022, doi: 10.1016/j.aej.2022.03.060.
A. Danielsen, R. Brøvig, K. K. Bøhler, G. S. Câmara, M. R. Haugen, E. Jacobsen, et al., “There’s more to timing than time: Investigating musical microrhythm across disciplines and cultures,” Music Perception: An Interdisciplinary Journal, vol. 41, no. 3, pp. 176-198, 2024, doi: 10.1525/mp.2024.41.3.176.
Y. Li and R. Sun, “Innovations of music and aesthetic education courses using intelligent technologies,” Education and Information Technologies, vol. 28, no. 10, pp. 13665-13688, 2023, doi: 10.1007/s10639-023-11624-9.
K. Tsioutas and G. Xylomenos, “On the impact of audio characteristics to the quality of musicians’ experience in network music performance,” J Audio Eng Soc, vol. 69, no. 12, pp. 914-923, 2021.
A. Stevenson, “The performance of machine aesthetics: acoustic reimagining of electronic music,” Popular Music, vol. 42, no. 4, pp. 438-456, 2023, doi: 10.1017/S026114302400014X.
X. Song and P. Phokha, “Applications of Artificial Intelligence-Assisted Computing In Piano Education,” Journal of Ecohumanism, vol. 3, no. 7, pp. 1648-1659, 2024.
M. Clayton, S. Tarsitani, R. Jankowsky, L. Jure, L. Leante, R. Polak, et al., “The interpersonal entrainment in music performance data collection,” Empirical Musicology Review, vol. 16, no. 1, pp. 65-84, 2021, doi: 10.18061/emr.v16i1.7555.