Process Modeling and Emotional Response Simulation of Musical Tension Evolution
Main Article Content
Abstract
Modeling the dynamic relationship between musical tension and emotional response remains challenging because existing approaches typically treat tension prediction and emotion recognition as independent tasks. This study proposes a process-oriented computational framework that integrates cognitive expectation with continuous emotional response simulation through hierarchical signal modeling and dynamic coupling mechanisms. A twelve-dimensional tension descriptor is extracted from symbolic scores and audio signals, while a Transformer-XL-enhanced IDyOM model estimates note-level surprise to characterize expectation violations. The resulting features are recursively encoded by a gated recurrent unit with expectation-aware modulation to generate continuous tension trajectories. Furthermore, a conditional variational autoencoder models listener-specific cognitive sensitivity, and a differential-equation-based coupling mechanism links tension evolution with arousal and valence dynamics to establish a closed-loop perception–expectation–response simulator. Experimental results demonstrate a peak localization error of 0.42 s, an arousal tracking RMSE of 0.152, a valence tracking RMSE of 0.176, and statistically significant causal effects of expectation modulation on simulation accuracy. The proposed framework provides an effective strategy for temporal signal fusion, dynamic state estimation, and adaptive response modeling, offering valuable references for intelligent audio signal processing, multimodal information perception, and computational sensing systems in electromagnetic signal analysis and next-generation interactive media applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
P. Kern, M. Heilbron, F. P. de Lange, and E. Spaak, “Cortical activity during naturalistic music listening reflects short-range predictions based on long-term experience,” elife, vol. 11, no. 1, Art. no. e80935, 2022, doi: 10.7554/eLife.80935.
X. Wang, L. Wang, and L. Xie, “Comparison and analysis of acoustic features of Western and Chinese classical music emotion recognition based on VA model,” Applied Sciences, vol. 12, no. 12, pp. 5787-5799, 2022, doi: 10.3390/app12125787.
V. K. M. Cheung, P. M. C. Harrison, S. Koelsch, M. T. Pearce, A. D. Friederici, and L. Meyer, “Cognitive and sensory expectations independently shape musical expectancy and pleasure,” Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 379, no. 1895, pp. 1-16, 2024, doi: 10.31234/osf.io/z76hg_v1.
N. Singer, N. Jacoby, T. Hendler, and R. Granot, “Feeling the beat: Temporal predictability is associated with ongoing changes in music-induced pleasantness,” Journal of Cognition, vol. 6, no. 1, pp. 34-51, 2023, doi: 10.5334/joc.286.
N. Zhang, L. Sun, Q. Wu, and Y. Yang, “Tension experience induced by tonal and melodic shift at music phrase boundaries,” Scientific Reports, vol. 12, no. 1, pp. 8304-8315, 2022, doi: 10.1038/s41598-022-11949-4.
K. Basi´nski, D. R. Quiroga-Martinez, and P. Vuust, “Temporal hierarchies in the predictive processing of melody-From pure tones to songs,” Neuroscience & Biobehavioral Reviews, vol. 145, no. 1, pp. 105007-105023, 2023, doi: 10.31234/osf.io/nua2j.
A. S. Turrell, A. R. Halpern, K. Bannister, D. Chai-Wi-Ting, and A. H. Javadi, “Building the anticipation: How variation in tension mediates emotions in music,” Music Perception: An Interdisciplinary Journal, vol. 42, no. 3, pp. 256-268, 2025, doi: 10.1525/mp.2024.aa004.
Y. Shang, Q. Peng, Z. Wu, and Y. Liu, “Music-induced emotion flow modeling by ENMI Network,” PLoS One, vol. 19, no. 10, Art. no. e0297712, 2024, doi: 10.1371/journal.pone.0297712.
S. You, L. Sun, and Y. Yang, “The effects of contextual certainty on tension induction and resolution,” Cognitive Neurodynamics, vol. 17, no. 1, pp. 191-201, 2023, doi: 10.1007/s11571-022-09810-5.
T. Daikoku, M. Tanaka, and S. Yamawaki, “Bodily maps of uncertainty and surprise in musical chord progression and the underlying emotional response,” Iscience, vol. 27, no. 4, pp. 109498-109510, 2024, doi: 10.1016/j.isci.2024.109498.
A. V. Barchet, J. M. Rimmele, and C. Pelofi, “TenseMusic: An automatic prediction model for musical tension,” Plos one, vol. 19, no. 1, Art. no. e0296385, 2024, doi: 10.31234/osf.io/xck3w.
G. Marion, F. Gao, B. P. Gold, G. M. Di Liberto, and S. Shamma, “IDyOMpy: A new Python-based model for the statistical analysis of musical expectations,” Journal of Neuroscience Methods, vol. 415, no. 1, Art. no. 110347, 2025, doi: 10.1016/j.jneumeth.2024.110347.
X. Guan, Z. Ren, and C. Pelofi, “py2lispIDyOM: A Python package for the information dynamics of music (IDyOM) model,” Journal of Open Source Software, vol. 7, no. 79, pp. 4738-4753, 2022, doi: 10.21105/joss.04738.
R. Orjesek, R. Jarina, and M. Chmulik, “End-to-end music emotion variation detection using iteratively reconstructed deep features,” Multimedia Tools and Applications, vol. 81, no. 4, pp. 5017-5031, 2022, doi: 10.1007/s11042-021-11584-7.
P. Bodily and D. Ventura, “Steerable music generation which satisfies long-range dependency constraints,” Transactions of the International Society for Music Information Retrieval, vol. 5, no. 1, pp. 71-86, 2022, doi: 10.5334/tismir.97.
M. R. Bjare, S. Lattner, and G. Widmer, “Differentiable short-term models for efficient online learning and prediction in monophonic music,” Transactions of the International Society for Music Information Retrieval, vol. 5, no. 1, pp. 190-206, 2022, doi: 10.5334/tismir.123.
M. Alfaro-Contreras, J. J. Valero-Mas, J. M. Iñesta, and J. Calvo-Zaragoza, “Late multimodal fusion for image and audio music transcription,” Expert Systems with Applications, vol. 216, no. 1, Art. no. 119491, 2023, doi: 10.1016/j.eswa.2022.119491.
W. Yu, “Music source feature extraction based on improved attention mechanism and phase feature,” Systems and Soft Computing, vol. 6, no. 1, Art. no. 200149, 2024, doi: 10.1016/j.sasc.2024.200149.
M. Navarro-Cáceres, M. Caetano, G. Bernardes, M. Sánchez-Barba, and J. Merchán Sánchez-Jara, “A computational model of tonal tension profile of chord progressions in the tonal interval space,” Entropy, vol. 22, no. 11, pp. 1291-1302, 2020, doi: 10.3390/e22111291.
F. Zhu, C. Wu, Q. Huang, N. Zhu, and T. Leng, “Rhythm-Based Attention Analysis: A Comprehensive Model for Music Hierarchy,” Applied Sciences, vol. 15, no. 11, pp. 6139-6152, 2025, doi: 10.3390/app15116139.
M. S. M. Mendjel, S. Ghazi, A. Dib, and H. Seridi, “A new audio approach based on user preferences analysis to enhance music recommendations,” Revue d’Intelligence Artificielle, vol. 37, no. 5, pp. 1341-1349, 2023, doi: 10.18280/ria.370527.
N. He and S. Ferguson, “Music emotion recognition based on segment-level two-stage learning,” International Journal of Multimedia Information Retrieval, vol. 11, no. 3, pp. 383-394, 2022, doi: 10.1007/s13735-022-00230-z.
A. J. Milne, E. A. Smit, H. S. Sarvasy, and R. T. Dean, “Evidence for a universal association of auditory roughness with musical stability,” PLoS One, vol. 18, no. 9, Art. no. e0291642, 2023, doi: 10.1371/journal.pone.0291642.
R. Marjieh, P. M. C. Harrison, H. Lee, F. Deligiannaki, and N. Jacoby, “Timbral effects on consonance disentangle psychoacoustic mechanisms and suggest perceptual origins for musical scales,” Nature Communications, vol. 15, no. 1, pp. 1482-1501, 2024, doi: 10.1038/s41467-024-45812-z.
V. Etxebarria, “Dissonance, Sound Spectrum and Musical Scale for Ancient Idiophones and Aerophones,” Journal of Mathematics and Music, vol. 1, no. 1, pp. 1-15, 2025, doi: 10.1080/17459737.2025.2560922.
J. Min, Z. Gao, and L. Wang, “Application and research of music generation system based on cvae and Transformer-XL in video background music,” IEEE Transactions on Industrial Informatics, vol. 21, no. 2, pp. 1409-1418, 2024, doi: 10.1109/TII.2024.3477561.
J. Liang, “Harmonizing minds and machines: survey on transformative power of machine learning in music,” Frontiers in Neurorobotics, vol. 17, no. 1, Art. no. 1267561, 2023, doi: 10.3389/fnbot.2023.1267561.
Y. Liang, H. Abudukelimu, J. Chen, A. Abulizi, and W. Guo, “MAML-XL: a symbolic music generation method based on meta-learning and Transformer-XL: Y,” Liang et al. Multimedia Systems, vol. 31, no. 3, pp. 206-225, 2025, doi: 10.1007/s00530-025-01803-8.
Z. Li, Q. Huang, X. Yang, Q. Chen, and L. Zhang, “Automatic composition system based on transformer-xl,” Applied Sciences, vol. 14, no. 13, pp. 5765-5776, 2024, doi: 10.3390/app14135765.
Y. Zhang, M. Li, and S. Pan, “Deep learning-based emotion recognition algorithms in music performance,” Scalable Computing: Practice and Experience, vol. 25, no. 6, pp. 4712-4719, 2024, doi: 10.12694/scpe.v25i6.3286.
N. B. Maimon, D. Lamy, and Z. Eitan, “Do Picardy thirds smile? Tonal hierarchy and tonal valence: Explicit and implicit measures,” Music Perception: An Interdisciplinary Journal, vol. 39, no. 5, pp. 443-467, 2022, doi: 10.1525/mp.2022.39.5.443.
S. Ji and X. Yang, “EmoMusicTV: Emotion-conditioned symbolic music generation with hierarchical transformer VAE,” IEEE Transactions on Multimedia, vol. 26, no. 1, pp. 1076-1088, 2023, doi: 10.1109/tmm.2023.3276177.
L. Comanducci, D. Gioiosa, M. Zanoni, F. Antonacci, and A. Sarti, “Variational Autoencoders for chord sequence generation conditioned on Western harmonic music complexity,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2023, no. 1, pp. 24-35, 2023, doi: 10.1186/s13636-023-00288-5.
Y. Xin, “MusicEmo: transformer-based intelligent approach towards music emotion generation and recognition,” Journal of Ambient Intelligence and Humanized Computing, vol. 15, no. 8, pp. 3107-3117, 2024, doi: 10.1007/s12652-024-04811-0.
S. L. Wu and Y. H. Yang, “MuseMorphose: Full-song and fine-grained piano music style transfer with one transformer VAE,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, no. 1, pp. 1953-1967, 2023, doi: 10.1109/TASLP.2023.3270726.