CLIP-ViL with Contrastive Learning for Automated Semantic Labeling of Digital Art

Main Article Content

Q. Fu
X. He

Abstract

With the rapid development of digital media technology, the number of digital artworks and intricate textile designs continues to grow, creating demand for efficient and accurate semantic annotation methods for management, retrieval, and dissemination. Existing digital art annotation methods suffer from low semantic alignment accuracy, insufficient generalization, and weak adaptability to artistic style diversity. To address these issues, this paper proposes an automated semantic labeling method for digital art based on the CLIP-ViL model combined with contrastive learning. The multimodal feature extraction framework of CLIP-ViL is first optimized to enhance the extraction of artistic style and semantic concept features. A cross-modal contrastive learning module is then designed to achieve fine-grained alignment between visual features and semantic labels. Finally, a multi-scale feature fusion strategy is introduced to improve recognition of multi-level semantic information. Experiments on ArtEmis, WikiArt, and DigitalArt-1M show that the method achieves average annotation accuracy of 92.3%, exceeding traditional methods by 18.7 percentage points and standalone CLIP by 9.5 percentage points. Recall, F1-score, and mAP reach 89.6%, 90.9%, and 91.8%, respectively.

Downloads

Download data is not yet available.

Article Details

How to Cite
Fu, Q., & He, X. (2026). CLIP-ViL with Contrastive Learning for Automated Semantic Labeling of Digital Art. Advanced Electromagnetics, 15(3), 5611–5620. https://doi.org/10.7716/aem.v15i3.3612
Section
Research Articles

References

A. Vaswani, N. Shazeer, and N. Parmar, “Attention is all you need,” Proceedings of the 31st International Conference on Neural Information Processing Systems; 4-9 December 2017; Long Beach, California, USA, pp. p. 5998-6008, 2017.

J. Li, L. Chen, and Y. Wang, “Application of CLIP-ViL in cross-modal semantic tasks,” Journal of Computational Design and Engineering, vol. 10, no. 2, pp. 567-582, 2023.

Y. Zhang, X. Liu, and W. Chen, “Research on intelligent management of digital art resources based on semantic annotation,” Journal of Cultural Heritage, vol. 65, pp. 234-245, 2023.

J. Li, L. Chen, and Y. Wang, “Automatic semantic labeling of digital art: A review,” Pattern Recognition, vol. 132, Art. no. 108876, 2022.

L. Wang, Y. Zhang, and J. Li, “Challenges and solutions of digital art semantic annotation,” Journal of Visual Communication and Image Representation, vol. 91, Art. no. 103678, 2023.

S. Chen, Y. Liu, and J. Zhang, “Multi-modal semantic annotation for digital artworks,” Multimedia Tools and Applications, vol. 81, no. 24, pp. 34567-34589, 2022.

A. Radford, K. Narasimhan, and T. Salimans, “Improving language understanding by generative pre-training,” 2018. arXiv preprint arXiv:1801.06146.

A. Dosovitskiy, L. Beyer, and A. Kolesnikov, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2020. arXiv preprint arXiv:2010.11929.

Z. Zhang, Y. Li, and C. Liu, “Semantic annotation of digital art: Problems and countermeasures,” Journal of the Association for Information Science and Technology, vol. 73, no. 8, pp. 1098-1112, 2022.

M. Smith, A. Johnson, and R. Miller, “Digital art classification using convolutional neural networks,” Pattern Recognition Letters, vol. 158, pp. 123-129, 2022.

A. Brown, B. Davis, and C. Wilson, “Transformer-based semantic labeling for digital paintings,” Computer Vision and Image Understanding, vol. 231, Art. no. 103456, 2023.

Sasse J, Fluck J.An Annotation Workbench for Semantic Annotation of Data Collection Instruments[J].Studies in health technology and informatics, 2023, 302:, 108-112, doi: 10.3233/SHTI230074.

View Article

Y. Liu, J. Li, and H. Zhang, “Cross-attention fusion for multimodal learning,” Neural Networks, vol. 158, pp. 234-245, 2023.

A. Radford, J. W. Kim, and C. Hallacy, “Learning transferable visual models from natural language supervision,” 2021. arXiv preprint arXiv:2103.00020.

Y. Cao, X. Lin, and L. Li, “ViT-GPT2: Hierarchical vision-language pre-training for image captioning and VQA,” 2021. arXiv preprint arXiv:2102.03334.

D. Borth, R. Ji, and T. Chen, “Large-scale visual sentiment ontology and detectors,” Proceedings of the International Conference on Multimedia; 2-5 December 2013; Luleå, Sweden, pp. p.223-232, 2013, doi: 10.1145/2502081.2502282.

View Article

T. L. Saaty, “The analytic hierarchy process: Planning, priority setting, resource allocation,” New York, NY, USA: McGraw-Hill; 1980.

T. Chen, S. Kornblith, and M. Norouzi, “A simple framework for contrastive learning of visual representations,” 2020. arXiv preprint arXiv:2002.05709.

K. He, H. Fan, and Y. Wu, “Momentum contrast for unsupervised visual representation learning,” 2019. arXiv preprint arXiv:1911.05722, doi: 10.1109/CVPR42600.2020.00975.

View Article

J. F. Hair, W. C. Black, and B. J. Babin, “Multivariate data analysis,” Harlow, UK: Pearson Education Limited; 2018.

D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” Proceedings of the 3rd International Conference on Learning Representations; 7-9 May 2015; San Diego, CA, USA; 2015.

T. Baltrušaitis, C. Ahuja, and P. Morency L, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443, 2019, doi: 10.1109/TPAMI.2018.2798607.

View Article

A. Zadeh, M. Chen, and S. Poria, “Multimodal fusion: A survey of methods, challenges, and applications,” IEEE Transactions on Knowledge and Data Engineering, vol. 34, no. 10, pp. 4712-4733, 2022.

H. Liu, W. Chen, and J. Zhang, “Multimodal fusion for semantic annotation,” Journal of Data and Information Quality, vol. 15, no. 2, pp. 1-20, 2023.

X. Wang, R. Girshick, and A. Gupta, “Non-local neural networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 18-22 June 2018; Salt Lake City, UT, USA, pp. p.7794-7803, 2018, doi: 10.1109/CVPR.2018.00813.

View Article

R. Hadsell, S. Chopra, and L. C. Yann, “Dimensionality reduction by learning an invariant mapping,” Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition; 17-22 June 2006; New York, NY, USA, pp. Volume 2. p.1735-1742, 2006, doi: 10.1109/CVPR.2006.100.

View Article

X. Chen, H. Fan, and R. Girshick, “Contrastive learning for cross-modal retrieval,” Computer Vision and Image Understanding, vol. 225, Art. no. 103867, 2022.

P. Lopes, J. Gama, and A. Veloso, “Semantics of digital art: A framework for analysis,” Journal of Visual Art Practice, vol. 21, no. 3, pp. 189-205, 2022.

F. Castro, M. Silva, and P. Costa, “Style and semantics in digital art classification,” Journal of New Media, vol. 18, no. 2, pp. 345-362, 2023.

J. Liu, H. Zhang, and L. Chen, “Semantic analysis of digital artworks,” Journal of Art Theory and Practice, vol. 14, no. 1, pp. 78-92, 2022.

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > >> 

You may also start an advanced similarity search for this article.