Research on Optimization of Personalized Recommendation Strategies for Short Videos
Main Article Content
Abstract
Personalized recommendation in short-video platforms requires accurate modeling of user interests from heterogeneous content signals while maintaining diversity, latency, and privacy constraints. To address data noise, recommendation homogenization, delayed interest capture, and ethical risks, this study proposes a multimodal recommendation optimization framework consisting of data governance, dynamic interest modeling, and compliance constraints. User behavior, video text, visual keyframes, audio features, and contextual information are collected through front-end logging and cleaned through anomaly filtering, missing-value processing, and differential privacy. CLIP, Swin Transformer, BERT, and MFCC features are integrated to form unified multimodal content representations. A three-stage recall–coarse-ranking–fine-ranking architecture is then developed, combining a dual-tower model, LightGBM, Wide&Deep, attention mechanism, dynamic interest decay, and PLE-based multi-task learning. Experiments on 1 million users and 5 million videos show that the optimized strategy improves click-through rate by 18. 3%, completion rate by 14.5%, content-type coverage by 22.1%, and average daily usage time by 15.7%, while reducing repeated recommendation and privacy-leakage risk. The framework supports multimodal signal processing, adaptive information delivery, and intelligent communication-service optimization.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
I. A. Monir, M. W. Fakhr, and N. El-Bendary, “Multimodal deep learning model for human handover classification,” Bulletin of Electrical Engineering and Informatics, vol. 11, no. 2, pp. 974-985, 2022, doi: 10.11591/eei.v11i2.3690.
N. Xie, Y. Liang, Z. Luo, et al., “Early prediction of adverse outcomes in liver cirrhosis using a CT-based multimodal deep learning model,” Abdominal Radiology, vol. 51, no. 1, pp. 137-150, 2026, doi: 10.1007/s00261-025-05045-0.
J. L. Koyner, J. Martin, K. A. Carey, et al., “Multicenter Development and Validation of a Multimodal Deep Learning Model to Predict Moderate to Severe AKI,” Clinical Journal of the American Society of Nephrology, vol. 20, no. 6, pp. 766-778, 2025, doi: 10.2215/CJN.0000000695.
M. Verlyck, D. Zhao, E. Ferdian, et al., “A Multimodal Deep Learning Model for Non-Invasive Detection of Elevated Left Ventricular End-Diastolic Pressure,” Heart, Lung & Circulation, vol. 34, no. Sup2, pp. S53-S54, 2025, doi: 10.1016/j.hlc.2025.04.063.
A. Bouatmane, A. Daaif, A. Bousselham, et al., “A Multimodal Deep Learning Model Integrating CNN and Trans former for Predicting Chemotherapy-Induced Cardiotoxicity,” IEEE Access, vol. 13, pp. 57568-57588, 2025, doi: 10.1109/ACCESS.2025.3556700.
Y. Lei, B. Feng, et al., “Predicting microvascular invasion in hepatocellular carcinoma with a CT-and MRI-based multi modal deep learning model,” Abdominal Radiology, vol. 49, no. 5, pp. 1397-1410, 2024, doi: 10.1007/s00261-024-04202-1.
G. Tian, “Multimodal Co-Ranking of Music Sentiment Analysis with the Deep Learning Model,” Journal of Electrical Systems, vol. 20, no. 6s, pp. 1446-1456, 2024, doi: 10.52783/jes.3059.
C. Yuan, Q. Shi, X. Huang, et al., “Multimodal deep learning model on interim [18F]FDG PET/CT for predicting primary treatment failure in diffuse large B-cell lymphoma,” European Radiology, vol. 33, no. 1, pp. 77-88, 2023, doi: 10.1007/s00330-022-09031-8.
W. H. Hui, W. H. Chiu, L. H. Yu, et al., “Reply: Technical considerations in the development of a multimodal deep learning model for predicting hepatocellular carcinoma outcomes,” Hepatology, vol. 83, no. 2, pp. E87-E88, 2025, doi: 10.1097/HEP.0000000000001540.
S. Athreya, A. Melehy, S. S. A. Suthahar, et al., “Combining Ultrasound Imaging and Molecular Testing in a Multimodal Deep Learning Model for Risk Stratification of Indeterminate Thyroid Nodules,” Thyroid, vol. 35, no. 5, pp. 590-594, 2025, doi: 10.1089/thy.2024.0584.
B. M. Faltas, Z. Bai, M. Osman, et al., “Predicting clinical outcomes in the S1314-COXEN trial using a multimodal deep learning model integrating histopathology, cell types, and gene expression,” Journal of Clinical Oncology, vol. 42, no. 4_suppl, pp. 533-533, 2024, doi: 10.1200/JCO.2024.42.4_suppl.533.
H. S. Jung, E. J. Lee, D. I. Chang, et al., “A Multimodal Ensemble Deep Learning Model for Functional Outcome Prognosis of Stroke Patients,” Journal of Stroke, vol. 26, no. 2, pp. 312-320, 2024, doi: 10.5853/jos.2023.03426.
Y. Zhang, G. Qin, S. Yang, et al., “Sea Surface Height Inversion Model Based on Multimodal Deep Learning for the Fusion of Heterogeneous FY-3E GNSS-R Data,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 11979-11995, 2025, doi: 10.1109/JSTARS.2025.3562741.