Movie Soundtrack Emotion Matching Algorithm Based on CLIP Model

Main Article Content

J. Y. Yang
X. Gao
W. X. Zhang

Abstract

To address the semantic distortion and temporal misalignment introduced by existing movie soundtrack matching methods that rely on textual intermediaries, this paper proposes an end-to-end audio-visual cross-modal temporal alignment algorithm based on the CLIP model. Considering that robust multimodal signal fusion and temporal synchronization are increasingly important for intelligent information processing in advanced electromagnetic sensing and communication environments, the proposed framework constructs a unified embedding space without textual bridging by extracting visual semantics through CLIP-ViT-B/32 and integrating AudioMAE with BiLSTM to capture deep temporal emotional characteristics of music. Musical representations are directly aligned with visual emotion prototypes to reduce the semantic gap, while a sliding-window temporal attention mechanism enables fine-grained dynamic soft alignment between scene transitions and musical emotional evolution. Furthermore, a multi-dimensional weighted scoring function incorporating semantic consistency, rhythm synchronization, and emotional intensity is designed, and global recommendation sequences are optimized under scene continuity constraints. Experimental results demonstrate that the proposed method achieves an emotion category accuracy of 76.3 ± 0.6%, an average cross-modal similarity of 0.732 ± 0.088, and an average temporal alignment deviation of only 1.24 ± 0.22 s. The framework significantly improves temporal synergy and emotional consistency between visual and audio modalities, providing an effective solution for cross-modal signal alignment and offering valuable methodological insights for multimodal information fusion and intelligent signal processing in electromagnetic wave perception and communication-oriented applications.

Downloads

Download data is not yet available.

Article Details

How to Cite
Yang, J. Y., Gao, X., & Zhang, W. X. (2026). Movie Soundtrack Emotion Matching Algorithm Based on CLIP Model. Advanced Electromagnetics, 15(3), 5287–5300. https://doi.org/10.7716/aem.v15i3.3583
Section
Research Articles

References

A. H. J. P. J. Permana and A. T. Wibowo, “Movie recommendation system based on synopsis using content-based filtering with TF-IDF and cosine similarity,” International Journal on Information and Communication Technology (IJoICT), vol. 9, no. 2, pp. 1-14, 2023, doi: 10.21108/ijoict.v9i2.

View Article

C. Jin, M. Lin, F. Wu, X. Wu, Y. Zhou, and J. Wang, “TVMTrailer: A Text-Video-Music AIGC Framework for Film Trailer Generation,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 55, no. 9, pp. 6000-6010, 2025, doi: 10.1109/TSMC.2025.3576988.

View Article

B. Millet, J. Chattah, and S. Ahn, “Soundtrack design: The impact of music on visual attention and affective responses,” Applied Ergonomics, vol. 93, Art. no. 103301, 2021, doi: 10.1016/j.apergo.2020.103301.

View Article

Y. Zhao, M. Yang, Y. Lin, X. Zhang, F. Shi, Wang, Z, et al., “AI-enabled text-to-music generation: A comprehensive review of methods, frameworks, and future directions,” Electronics, vol. 14, no. 6, pp. 1197, 2025, doi: 10.3390/electronics14061197.

View Article

J. Gan, “Music feature classification based on recurrent neural networks with channel attention mechanism,” Mobile Information Systems, vol. 2021, no. 1, Art. no. 7629994, 2021, doi: 10.1155/2021/7629994.

View Article

G. Tong, “Multimodal music emotion recognition method based on the combination of knowledge distillation and transfer learning,” Scientific Programming, vol. 2022, no. 1, Art. no. 2802573, 2022, doi: 10.1155/2022/2802573.

View Article

Y. R. Pandeya and J. Lee, “Deep learning-based late fusion of multimodal information for emotion classification of music video,” Multimedia Tools and Applications, vol. 80, no. 2, pp. 2887-2905, 2021, doi: 10.1007/s11042-020-08836-3.

View Article

J. Salas-Cáceres, J. Lorenzo-Navarro, D. Freire-Obregón, and M. Castrillón-Santana, “Multimodal emotion recognition based on a fusion of audiovisual information with temporal dynamics,” Multimedia Tools and Applications, vol. 84, pp. 27327-27343, 2024, doi: 10.1007/s11042-024-20227-6.

View Article

S. Moorthy and Y. K. Moon, “Hybrid Multi-Attention Network for Audio– Visual Emotion Recognition Through Multimodal Feature Fusion,” Mathematics, vol. 13, no. 7, pp. 1100, 2025, doi: 10.3390/math13071100.

View Article

Z. Su, Y. Feng, J. Liu, J. Peng, W. Jiang, and J. Liu, “An Audiovisual Correlation Matching Method Based on Fine-Grained Emotion and Feature Fusion,” Sensors, vol. 24, no. 17, pp. 5681, 2024, doi: 10.3390/s24175681.

View Article

Q. Shen, J. Xu, J. Mei, X. Wu, and D. Dong, “EmoStyle: Emotion-aware semantic image manipulation with audio guidance,” Applied Sciences, vol. 14, no. 8, pp. 3193, 2024, doi: 10.3390/app14083193.

View Article

J. Li, W. K. Wong, L. Jiang, X. Fang, S. Xie, and Y. Xu, “CKDH: CLIP-based knowledge distillation hashing for cross-modal retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 7, pp. 6530-6541, 2024, doi: 10.1109/TCSVT.2024.3350695.

View Article

J. Yang, L. Wang, Y. Qi, H. Chen, and J. Li, “Multimodal Information Fusion and Data Generation for Evaluation of Second Language Emotional Expression,” Applied Sciences, vol. 14, no. 19, pp. 9121, 2024, doi: 10.3390/app14199121.

View Article

G. Yang, K. Xu, X. Fang, and J. Zhang, “Video face forgery detection via facial motion-assisted capturing dense optical flow truncation,” The Visual Computer, vol. 39, no. 11, pp. 5589-5608, 2023, doi: 10.1007/s00371-022-02683-z.

View Article

T. K. T. Zizi, S. Ramli, M. Wook, and M. A. M. Shukran, “Optical flow-based algorithm analysis to detect human emotion from eye movement-image data,” Journal of Image and Graphics, vol. 11, no. 1, pp. 53-60, 2023, doi: 10.18178/joig.11.1.53-60.

View Article

Y. Shao and N. Guo, “Recognizing online video genres using ensemble deep convolutional learning for digital media service management,” Journal of Cloud Computing, vol. 13, no. 1, pp. 102, 2024, doi: 10.1186/s13677-024-00664-2.

View Article

R. Tan, M. Sun, and Y. Liang, “Transformer-based multi-level attention integration network for video saliency prediction,” Multimedia Tools and Applications, vol. 84, no. 13, pp. 11833-11854, 2025, doi: 10.1007/s11042-024-19404-4.

View Article

J. Min, Z. Gao, L. Wang, and A. Zhang, “Application research of short-time Fourier transform in music generation based on the parallel WaveGan system,” IEEE Transactions on Industrial Informatics, vol. 20, no. 9, pp. 10770-10778, 2024, doi: 10.1109/TII.2024.3397344.

View Article

C. Constantinescu and R. Brad, “An overview of sound features in time and frequency domain,” International Journal of Advanced Statistics and IT&C for Economics and Life Sciences, vol. 13, no. 1, pp. 45-58, 2023, doi: 10.2478/ijasitels-2023-0006.

View Article

S. Seo, C. Kim, and J. H. Kim, “Convolutional neural networks using log mel-spectrogram separation for audio event classification with unknown devices,” Journal of Web Engineering, vol. 21, no. 2, pp. 497-522, 2022, doi: 10.13052/jwe1540-9589.21216.

View Article

L. C. Reghunath and R. Rajan, “Transformer-based ensemble method for multiple predominant instruments recognition in polyphonic music,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2022, no. 1, pp. 11, 2022, doi: 10.1186/s13636-022-00245-8.

View Article

V. K. Singh, K. Sharma, and S. N. Sur, “Acoustic scene classification using dynamic time warping technique based on short time Fourier transform and discrete wavelet transforms,” Circuits, Systems, and Signal Processing, vol. 44, no. 3, pp. 1887-1913, 2025, doi: 10.1007/s00034-024-02895-9.

View Article

V. K. Singh, K. Sharma, and S. N. Sur, “Impact of various continuous wavelet transforms for acoustic scene classification with DCASE dataset,” Signal, Image and Video Processing, vol. 19, no. 6, pp. 440, 2025, doi: 10.1007/s11760-025-04031-9.

View Article

R. De Prisco, A. Guarino, D. Malandrino, and R. Zaccagnino, “Induced emotion-based music recommendation through reinforcement learning,” Applied Sciences, vol. 12, no. 21, Art. no. 11209, 2022, doi: 10.3390/app122111209.

View Article

Similar Articles

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > >> 

You may also start an advanced similarity search for this article.