CLIP-ViL with Contrastive Learning for Automated Semantic Labeling of Digital Art
Main Article Content
Abstract
With the rapid development of digital media technology, the number of digital artworks and intricate textile designs continues to grow, creating demand for efficient and accurate semantic annotation methods for management, retrieval, and dissemination. Existing digital art annotation methods suffer from low semantic alignment accuracy, insufficient generalization, and weak adaptability to artistic style diversity. To address these issues, this paper proposes an automated semantic labeling method for digital art based on the CLIP-ViL model combined with contrastive learning. The multimodal feature extraction framework of CLIP-ViL is first optimized to enhance the extraction of artistic style and semantic concept features. A cross-modal contrastive learning module is then designed to achieve fine-grained alignment between visual features and semantic labels. Finally, a multi-scale feature fusion strategy is introduced to improve recognition of multi-level semantic information. Experiments on ArtEmis, WikiArt, and DigitalArt-1M show that the method achieves average annotation accuracy of 92.3%, exceeding traditional methods by 18.7 percentage points and standalone CLIP by 9.5 percentage points. Recall, F1-score, and mAP reach 89.6%, 90.9%, and 91.8%, respectively.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
A. Vaswani, N. Shazeer, and N. Parmar, “Attention is all you need,” Proceedings of the 31st International Conference on Neural Information Processing Systems; 4-9 December 2017; Long Beach, California, USA, pp. p. 5998-6008, 2017.
J. Li, L. Chen, and Y. Wang, “Application of CLIP-ViL in cross-modal semantic tasks,” Journal of Computational Design and Engineering, vol. 10, no. 2, pp. 567-582, 2023.
Y. Zhang, X. Liu, and W. Chen, “Research on intelligent management of digital art resources based on semantic annotation,” Journal of Cultural Heritage, vol. 65, pp. 234-245, 2023.
J. Li, L. Chen, and Y. Wang, “Automatic semantic labeling of digital art: A review,” Pattern Recognition, vol. 132, Art. no. 108876, 2022.
L. Wang, Y. Zhang, and J. Li, “Challenges and solutions of digital art semantic annotation,” Journal of Visual Communication and Image Representation, vol. 91, Art. no. 103678, 2023.
S. Chen, Y. Liu, and J. Zhang, “Multi-modal semantic annotation for digital artworks,” Multimedia Tools and Applications, vol. 81, no. 24, pp. 34567-34589, 2022.
A. Radford, K. Narasimhan, and T. Salimans, “Improving language understanding by generative pre-training,” 2018. arXiv preprint arXiv:1801.06146.
A. Dosovitskiy, L. Beyer, and A. Kolesnikov, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2020. arXiv preprint arXiv:2010.11929.
Z. Zhang, Y. Li, and C. Liu, “Semantic annotation of digital art: Problems and countermeasures,” Journal of the Association for Information Science and Technology, vol. 73, no. 8, pp. 1098-1112, 2022.
M. Smith, A. Johnson, and R. Miller, “Digital art classification using convolutional neural networks,” Pattern Recognition Letters, vol. 158, pp. 123-129, 2022.
A. Brown, B. Davis, and C. Wilson, “Transformer-based semantic labeling for digital paintings,” Computer Vision and Image Understanding, vol. 231, Art. no. 103456, 2023.
Sasse J, Fluck J.An Annotation Workbench for Semantic Annotation of Data Collection Instruments[J].Studies in health technology and informatics, 2023, 302:, 108-112, doi: 10.3233/SHTI230074.
Y. Liu, J. Li, and H. Zhang, “Cross-attention fusion for multimodal learning,” Neural Networks, vol. 158, pp. 234-245, 2023.
A. Radford, J. W. Kim, and C. Hallacy, “Learning transferable visual models from natural language supervision,” 2021. arXiv preprint arXiv:2103.00020.
Y. Cao, X. Lin, and L. Li, “ViT-GPT2: Hierarchical vision-language pre-training for image captioning and VQA,” 2021. arXiv preprint arXiv:2102.03334.
D. Borth, R. Ji, and T. Chen, “Large-scale visual sentiment ontology and detectors,” Proceedings of the International Conference on Multimedia; 2-5 December 2013; Luleå, Sweden, pp. p.223-232, 2013, doi: 10.1145/2502081.2502282.
T. L. Saaty, “The analytic hierarchy process: Planning, priority setting, resource allocation,” New York, NY, USA: McGraw-Hill; 1980.
T. Chen, S. Kornblith, and M. Norouzi, “A simple framework for contrastive learning of visual representations,” 2020. arXiv preprint arXiv:2002.05709.
K. He, H. Fan, and Y. Wu, “Momentum contrast for unsupervised visual representation learning,” 2019. arXiv preprint arXiv:1911.05722, doi: 10.1109/CVPR42600.2020.00975.
J. F. Hair, W. C. Black, and B. J. Babin, “Multivariate data analysis,” Harlow, UK: Pearson Education Limited; 2018.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” Proceedings of the 3rd International Conference on Learning Representations; 7-9 May 2015; San Diego, CA, USA; 2015.
T. Baltrušaitis, C. Ahuja, and P. Morency L, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443, 2019, doi: 10.1109/TPAMI.2018.2798607.
A. Zadeh, M. Chen, and S. Poria, “Multimodal fusion: A survey of methods, challenges, and applications,” IEEE Transactions on Knowledge and Data Engineering, vol. 34, no. 10, pp. 4712-4733, 2022.
H. Liu, W. Chen, and J. Zhang, “Multimodal fusion for semantic annotation,” Journal of Data and Information Quality, vol. 15, no. 2, pp. 1-20, 2023.
X. Wang, R. Girshick, and A. Gupta, “Non-local neural networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 18-22 June 2018; Salt Lake City, UT, USA, pp. p.7794-7803, 2018, doi: 10.1109/CVPR.2018.00813.
R. Hadsell, S. Chopra, and L. C. Yann, “Dimensionality reduction by learning an invariant mapping,” Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition; 17-22 June 2006; New York, NY, USA, pp. Volume 2. p.1735-1742, 2006, doi: 10.1109/CVPR.2006.100.
X. Chen, H. Fan, and R. Girshick, “Contrastive learning for cross-modal retrieval,” Computer Vision and Image Understanding, vol. 225, Art. no. 103867, 2022.
P. Lopes, J. Gama, and A. Veloso, “Semantics of digital art: A framework for analysis,” Journal of Visual Art Practice, vol. 21, no. 3, pp. 189-205, 2022.
F. Castro, M. Silva, and P. Costa, “Style and semantics in digital art classification,” Journal of New Media, vol. 18, no. 2, pp. 345-362, 2023.
J. Liu, H. Zhang, and L. Chen, “Semantic analysis of digital artworks,” Journal of Art Theory and Practice, vol. 14, no. 1, pp. 78-92, 2022.