Combining CRNN Modeling to Dynamically Align the Relationship Between Music Rhythm and Plot Twists in Different Film Genres
Main Article Content
Abstract
This paper addresses the challenge of quantitatively modeling the dynamic alignment between musical rhythm and plot turning points in films by proposing V2A-AlignNet, a genre-aware cross-modal deep learning framework. As intelligent multimedia perception increasingly relies on advanced signal processing and multimodal information fusion techniques that are conceptually relevant to electromagnetic sensing and communication systems, accurate temporal alignment has become an important research topic. The proposed model adopts a dual-stream architecture integrating VideoMAE for long-range spatiotemporal video representation and a CRNN for extracting both local and global rhythmic characteristics from Mel spectrograms. A cross-modal attention module constructs a shared semantic space to generate alignment saliency sequences and similarity matrices, while a genre-conditioning mechanism enables adaptive modeling for different film categories. Experiments conducted on 120 films spanning six genres (800 clips) demonstrate that V2A-AlignNet achieves superior performance in turning-point detection, alignment pattern classification, and genre recognition compared with representative baseline methods. Ablation studies further verify the effectiveness of each component, and visualization results reveal distinctive genre-specific alignment behaviors. The proposed framework provides a computational basis for audio-visual temporal relationship analysis and offers valuable insights for multimodal signal interpretation and intelligent information processing in advanced electromagnetic sensing and communication-related applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
N. Elyamany, “A chronotopic approach to identity performance in musical numbers: a choreo-musical case study of ‘rewrite the stars’ and ‘this is me’,” Visual Communication, vol. 22, no. 2, pp. 278-296, 2023, doi: 10.1177/1470357220974069.
A. K. Herget, “On music’s potential to convey meaning in film: A systematic review of empirical evidence,” Psychology of Music, vol. 49, no. 1, pp. 21-49, 2021, doi: 10.1177/0305735619835019.
T. Behrouzi, R. Toosi, and M. A. Akhaee, “Multimodal movie genre classification using recurrent neural network,” Multimedia Tools and Applications, vol. 82, no. 4, pp. 5763-5784, 2023, doi: 10.1007/s11042-022-13418-6.
A. K. Herget, “Well-known and unknown music as an emotionalizing carrier of meaning in film,” Media Psychology, vol. 24, no. 3, pp. 385-412, 2021, doi: 10.1080/15213269.2020.1713164.
W. Q. Liu, M. X. Lin, H. B. Huang, C. Y. Ma, Y. Song, W. M. Dong, et al., “Emotion-Aware Music Driven Movie Montage,” Journal of Computer Science and Technology, vol. 38, no. 3, pp. 540-553, 2023, doi: 10.1007/s11390-023-3064-6.
M. J. Lucia-Mulas, P. Revuelta-Sanz, B. Ruiz-Mezcua, and I. Gonzalez-Carrasco, “Automatic music emotion classification model for movie soundtrack subtitling based on neuroscientific premises,” Applied Intelligence, vol. 53, no. 22, pp. 27096-27109, 2023, doi: 10.1007/s10489-023-04967-w.
S. G. Cunha and M. Lee, “Emotional Video to Audio Transformation Using Deep Recurrent Neural Networks and a Neuro-Fuzzy System,” Mathematical Problems in Engineering, vol. 2020, no. 1, 8478527, 2020, doi: 10.1155/2020/8478527.
Y. R. Pandeya, B. Bhattarai, and J. Lee, “Music video emotion classification using slow–fast audio–video network and unsupervised feature representation,” Scientific Reports, vol. 11, no. 1, 19834, 2021, doi: 10.1038/s41598-021-98856-2.
J. Zhang, X. Wen, and M. Whang, “Recognition of emotion according to the physical elements of the video,” Sensors, vol. 20, no. 3, 649, 2020, doi: 10.3390/s20030649.
A. Sharma, K. Sharma, and A. Kumar, “Real-time emotional health detection using fine-tuned transfer networks with multimodal fusion,” Neural computing and applications, vol. 35, no. 31, pp. 22935-22948, 2023, doi: 10.1007/s00521-022-06913-2.
J. Tian and Y. She, “A visual–audio-based emotion recognition system integrating dimensional analysis,” IEEE Transactions on Computational Social Systems, vol. 10, no. 6, pp. 3273-3282, 2022, doi: 10.1109/TCSS.2022.3200060.
T. W. Kim and K. C. Kwak, “Speech emotion recognition using deep learning transfer models and explainable techniques,” Applied Sciences, vol. 14, no. 4, 1553, 2024, doi: 10.3390/app14041553.
S. Deng, L. Wu, G. Shi, L. H. Xing, W. J. Hu, and H. Zhang, “Simple but powerful, a language-supervised method for image emotion classification,” IEEE Transactions on Affective Computing, vol. 14, no. 4, pp. 3317-3331, 2022, doi: 10.1109/taffc.2022.3225049.
G. Shi, S. Deng, B. Wang, C. Feng, Y. Zhuang, and X. M. Wang, “One for all: A unified generative framework for image emotion classification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 8, pp. 7057-7068, 2023, doi: 10.1109/TCSVT.2023.3341840.
E. A. Iyilikci, A. Demirel, F. Isik, and O. Iyilkci, “How do sound and color features affect self-report emotional experience in response to film clips?,” Current Psychology, vol. 43, no. 11, pp. 10185-10216, 2024, doi: 10.1007/s12144-023-05127-6.
F. Liu, D. L. Chen, R. Z. Zhou, S. Yang, and F. Xu, “Self-supervised music motion synchronization learning for music-driven conducting motion generation,” Journal of Computer Science and Technology, vol. 37, no. 3, pp. 539-558, 2022, doi: 10.1007/s11390-022-2030-z.
S. Landry and M. Jeon, “Interactive sonification strategies for the motion and emotion of dance performances,” Journal on Multimodal User Interfaces, vol. 14, no. 2, pp. 167-186, 2020, doi: 10.1007/s12193-020-00321-3.
G. Bente, K. Kryston, N. T. Jahn, and R. Schmalzle, “Building blocks of suspense: subjective and physiological effects of narrative content and film music,” Humanities and Social Sciences Communications, vol. 9, no. 1, pp. 1-13, 2022, doi: 10.1057/s41599-022-01461-5.
E. M. Khan, M. S. Mukta, M. E. Ali, and J. Mahmud, “Predicting users’ movie preference and rating behavior from personality and values,” ACM Transactions on Interactive Intelligent Systems (TiiS), vol. 10, no. 3, pp. 1-25, 2020, doi: 10.1145/3338244.
J. D. Bradbury and R. E. Guadagno, “Documentary narrative visualization: Features and modes of documentary film in narrative visualization,” Information Visualization, vol. 19, no. 4, pp. 339-352, 2020, doi: 10.1177/1473871620925071.
B. N. Alkahtani, N. Alsraisri, and A. J. Almalki, “Achieving State-of-the-art Accuracy in Arabic Sign Language Recognition with VideoMAE and Data Augmentation,” Journal of Disability Research, vol. 4, no. 5, 20250661, 2025, doi: 10.57197/JDR-2025-0661.
A. Berroukham, M. Lahraichi, and K. Housni, “Self-Supervised Method for Risky Situation Detection in Road Traffic Sequences Using Video Masked Autoencoder,” International Journal of Advanced Computer Science & Applications, vol. 16, no. 6, pp. 559-559, 2025, doi: 10.14569/ijacsa.2025.0160654.
M. Suresha, S. Kuppa, and D. S. Raghukumar, “A study on deep learning spatiotemporal models and feature extraction techniques for video understanding,” International Journal of Multimedia Information Retrieval, vol. 9, no. 2, pp. 81-101, 2020, doi: 10.1007/s13735-019-00190-x.
M. S. Islam, S. Sultana, R. U. Kumar, and J. A. Mahmud, “A review on video classification with methods, findings, performance, challenges, limitations and future work,” Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, vol. 6, no. 2, pp. 47-57, 2020, doi: 10.26555/jiteki.v6i2.18978.
Y. Mao, G. Zhong, H. Wang, and K. Huang, “Music-CRN: An efficient content-based music classification and recommendation network,” Cognitive Computation, vol. 14, no. 6, pp. 2306-2316, 2022, doi: 10.1007/s12559-022-10039-x.
A. Bansal and N. K. Garg, “Robust technique for environmental sound classification using convolutional recurrent neural network,” Multimedia Tools and Applications, vol. 83, no. 18, pp. 54755-54772, 2024, doi: 10.1007/s11042-023-17066-2.
J. Y. Kwak and Y. J. Chung, “Audio Event Detection Based on Attention CRNN,” The Journal of the Korea institute of electronic communication sciences, vol. 15, no. 3, pp. 465-472, 2020, doi: 10.13067/JKIECS.2020.15.3.465.
M. P. Fatah, “CRNN Algorithm and MFCC Feature Extraction in Classifying Hijaiyah Letter Pronunciation: A Systematic Literature Review,” Khazanah Journal of Religion and Technology, vol. 3, no. 1, pp. 31-35, 2025, doi: 10.15575/kjrt.v3i1.1555.
S. Seo, C. Kim, and J. H. Kim, “Convolutional neural networks using log mel-spectrogram separation for audio event classification with unknown devices,” Journal of Web Engineering, vol. 21, no. 2, pp. 497-522, 2022, doi: 10.13052/jwe1540-9589.21216.
P. Rawat, M. Bajaj, S. Vats, and V. Sharma, “A comprehensive study based on MFCC and spectrogram for audio classification,” Journal of Information and Optimization Sciences, vol. 44, no. 6, pp. 1057-1074, 2023, doi: 10.47974/jios-1431.
N. Messina, G. Amato, A. Esuli, F. Falchi, C. Gennaro, S. Marchand-Maillet, et al., “Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders,” ACM Transactions on Multimedia Computing Communications, and Applications (TOMM), vol. 17, no. 4, pp. 1-23, 2021, doi: 10.1145/3451390.
X. Xu, T. Wang, Y. Yang, L. Zuo, F. Shen, and H. T. Shen, “Cross-modal attention with semantic consistence for image–text matching,” IEEE transactions on neural networks and learning systems, vol. 31, no. 12, pp. 5412-5425, 2020, doi: 10.1109/TNNLS.2020.2967597.
L. Li, T. Jin, W. Lin, H. Jiang, W. W. Pan, and J. Wang, “Multi-granularity relational attention network for audio-visual question answering,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 8, pp. 7080-7094, 2023, doi: 10.1109/TCSVT.2023.3264524.
L. Ye, M. Rochan, Z. Liu, X. Q. Zhang, and Y. Wang, “Referring segmentation in images and videos with cross-modal self-attention network,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 7, pp. 3719-3732, 2021, doi: 10.1109/TPAMI.2021.3054384.
P. Wei, F. He, L. Li, and J. Li, “Research on sound classification based on SVM,” Neural Computing and Applications, vol. 32, no. 6, pp. 1593-1607, 2020, doi: 10.1007/s00521-019-04182-0.
S. Accattoli, P. Sernani, N. Falcionelli, D. N. Mekuria, and A. Dragoni, “Violence detection in videos by combining 3D convolutional neural networks and support vector machines,” Applied Artificial Intelligence, vol. 34, no. 4, pp. 329-344, 2020, doi: 10.1080/08839514.2020.1723876.