Movie Soundtrack Emotion Matching Algorithm Based on CLIP Model
Main Article Content
Abstract
To address the semantic distortion and temporal misalignment introduced by existing movie soundtrack matching methods that rely on textual intermediaries, this paper proposes an end-to-end audio-visual cross-modal temporal alignment algorithm based on the CLIP model. Considering that robust multimodal signal fusion and temporal synchronization are increasingly important for intelligent information processing in advanced electromagnetic sensing and communication environments, the proposed framework constructs a unified embedding space without textual bridging by extracting visual semantics through CLIP-ViT-B/32 and integrating AudioMAE with BiLSTM to capture deep temporal emotional characteristics of music. Musical representations are directly aligned with visual emotion prototypes to reduce the semantic gap, while a sliding-window temporal attention mechanism enables fine-grained dynamic soft alignment between scene transitions and musical emotional evolution. Furthermore, a multi-dimensional weighted scoring function incorporating semantic consistency, rhythm synchronization, and emotional intensity is designed, and global recommendation sequences are optimized under scene continuity constraints. Experimental results demonstrate that the proposed method achieves an emotion category accuracy of 76.3 ± 0.6%, an average cross-modal similarity of 0.732 ± 0.088, and an average temporal alignment deviation of only 1.24 ± 0.22 s. The framework significantly improves temporal synergy and emotional consistency between visual and audio modalities, providing an effective solution for cross-modal signal alignment and offering valuable methodological insights for multimodal information fusion and intelligent signal processing in electromagnetic wave perception and communication-oriented applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
A. H. J. P. J. Permana and A. T. Wibowo, “Movie recommendation system based on synopsis using content-based filtering with TF-IDF and cosine similarity,” International Journal on Information and Communication Technology (IJoICT), vol. 9, no. 2, pp. 1-14, 2023, doi: 10.21108/ijoict.v9i2.
C. Jin, M. Lin, F. Wu, X. Wu, Y. Zhou, and J. Wang, “TVMTrailer: A Text-Video-Music AIGC Framework for Film Trailer Generation,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 55, no. 9, pp. 6000-6010, 2025, doi: 10.1109/TSMC.2025.3576988.
B. Millet, J. Chattah, and S. Ahn, “Soundtrack design: The impact of music on visual attention and affective responses,” Applied Ergonomics, vol. 93, Art. no. 103301, 2021, doi: 10.1016/j.apergo.2020.103301.
Y. Zhao, M. Yang, Y. Lin, X. Zhang, F. Shi, Wang, Z, et al., “AI-enabled text-to-music generation: A comprehensive review of methods, frameworks, and future directions,” Electronics, vol. 14, no. 6, pp. 1197, 2025, doi: 10.3390/electronics14061197.
J. Gan, “Music feature classification based on recurrent neural networks with channel attention mechanism,” Mobile Information Systems, vol. 2021, no. 1, Art. no. 7629994, 2021, doi: 10.1155/2021/7629994.
G. Tong, “Multimodal music emotion recognition method based on the combination of knowledge distillation and transfer learning,” Scientific Programming, vol. 2022, no. 1, Art. no. 2802573, 2022, doi: 10.1155/2022/2802573.
Y. R. Pandeya and J. Lee, “Deep learning-based late fusion of multimodal information for emotion classification of music video,” Multimedia Tools and Applications, vol. 80, no. 2, pp. 2887-2905, 2021, doi: 10.1007/s11042-020-08836-3.
J. Salas-Cáceres, J. Lorenzo-Navarro, D. Freire-Obregón, and M. Castrillón-Santana, “Multimodal emotion recognition based on a fusion of audiovisual information with temporal dynamics,” Multimedia Tools and Applications, vol. 84, pp. 27327-27343, 2024, doi: 10.1007/s11042-024-20227-6.
S. Moorthy and Y. K. Moon, “Hybrid Multi-Attention Network for Audio– Visual Emotion Recognition Through Multimodal Feature Fusion,” Mathematics, vol. 13, no. 7, pp. 1100, 2025, doi: 10.3390/math13071100.
Z. Su, Y. Feng, J. Liu, J. Peng, W. Jiang, and J. Liu, “An Audiovisual Correlation Matching Method Based on Fine-Grained Emotion and Feature Fusion,” Sensors, vol. 24, no. 17, pp. 5681, 2024, doi: 10.3390/s24175681.
Q. Shen, J. Xu, J. Mei, X. Wu, and D. Dong, “EmoStyle: Emotion-aware semantic image manipulation with audio guidance,” Applied Sciences, vol. 14, no. 8, pp. 3193, 2024, doi: 10.3390/app14083193.
J. Li, W. K. Wong, L. Jiang, X. Fang, S. Xie, and Y. Xu, “CKDH: CLIP-based knowledge distillation hashing for cross-modal retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 6530-6541, 2024, doi: 10.1109/TCSVT.2024.3350695.
J. Yang, L. Wang, Y. Qi, H. Chen, and J. Li, “Multimodal Information Fusion and Data Generation for Evaluation of Second Language Emotional Expression,” Applied Sciences, vol. 14, no. 19, pp. 9121, 2024, doi: 10.3390/app14199121.
G. Yang, K. Xu, X. Fang, and J. Zhang, “Video face forgery detection via facial motion-assisted capturing dense optical flow truncation,” The Visual Computer, vol. 39, no. 11, pp. 5589-5608, 2023, doi: 10.1007/s00371-022-02683-z.
T. K. T. Zizi, S. Ramli, M. Wook, and M. A. M. Shukran, “Optical flow-based algorithm analysis to detect human emotion from eye movement-image data,” Journal of Image and Graphics, vol. 11, no. 1, pp. 53-60, 2023, doi: 10.18178/joig.11.1.53-60.
Y. Shao and N. Guo, “Recognizing online video genres using ensemble deep convolutional learning for digital media service management,” Journal of Cloud Computing, vol. 13, no. 1, pp. 102, 2024, doi: 10.1186/s13677-024-00664-2.
R. Tan, M. Sun, and Y. Liang, “Transformer-based multi-level attention integration network for video saliency prediction,” Multimedia Tools and Applications, vol. 84, no. 13, pp. 11833-11854, 2025, doi: 10.1007/s11042-024-19404-4.
J. Min, Z. Gao, L. Wang, and A. Zhang, “Application research of short-time Fourier transform in music generation based on the parallel WaveGan system,” IEEE Transactions on Industrial Informatics, vol. 20, no. 9, pp. 10770-10778, 2024, doi: 10.1109/TII.2024.3397344.
C. Constantinescu and R. Brad, “An overview of sound features in time and frequency domain,” International Journal of Advanced Statistics and IT&C for Economics and Life Sciences, vol. 13, no. 1, pp. 45-58, 2023, doi: 10.2478/ijasitels-2023-0006.
S. Seo, C. Kim, and J. H. Kim, “Convolutional neural networks using log mel-spectrogram separation for audio event classification with unknown devices,” Journal of Web Engineering, vol. 21, no. 2, pp. 497-522, 2022, doi: 10.13052/jwe1540-9589.21216.
L. C. Reghunath and R. Rajan, “Transformer-based ensemble method for multiple predominant instruments recognition in polyphonic music,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2022, no. 1, pp. 11, 2022, doi: 10.1186/s13636-022-00245-8.
V. K. Singh, K. Sharma, and S. N. Sur, “Acoustic scene classification using dynamic time warping technique based on short time Fourier transform and discrete wavelet transforms,” Circuits, Systems, and Signal Processing, vol. 44, no. 3, pp. 1887-1913, 2025, doi: 10.1007/s00034-024-02895-9.
V. K. Singh, K. Sharma, and S. N. Sur, “Impact of various continuous wavelet transforms for acoustic scene classification with DCASE dataset,” Signal, Image and Video Processing, vol. 19, no. 6, pp. 440, 2025, doi: 10.1007/s11760-025-04031-9.
R. De Prisco, A. Guarino, D. Malandrino, and R. Zaccagnino, “Induced emotion-based music recommendation through reinforcement learning,” Applied Sciences, vol. 12, no. 21, Art. no. 11209, 2022, doi: 10.3390/app122111209.