Optimizing Cross-Modal Image–Text Matching Algorithms Based on Graph Neural Networks: Implications for Textile Applications

Main Article Content

K. Dong

Abstract

The challenge of image–text matching stems from differences in representation and structural semantics, particularly in complex scenarios such as digital textile and apparel analysis, resulting in a semantic gap that degrades matching performance. This paper proposes an optimized graph neural network (GNN)-based algorithm to improve cross-modal matching accuracy. Images and text are first represented as graph structures, where image regions and textual keywords are modeled as nodes and their relationships are encoded as edges. The GNN enables progressive information propagation and feature fusion across modalities, while an adaptive semantic alignment module dynamically adjusts propagation weights according to node-level semantic similarity. A multi-scale GNN architecture is further introduced to capture hierarchical semantic features and is optimized using a cross-modal contrastive loss that maximizes the similarity of correctly matched image–text pairs. Such structured multimodal semantic modeling also provides methodological insights for intelligent information fusion and heterogeneous data interpretation in advanced electromagnetic sensing and communication systems. Experimental results demonstrate Recall@1 scores of 70.1% and 60.5% on the MSCOCO and Flickr30k datasets, respectively, with a mean rank of 3.1 and mutual information of 0.75. The proposed method effectively narrows the semantic gap and significantly enhances cross-modal matching accuracy, demonstrating substantial potential for intelligent retrieval and analysis in the textile industry.

Downloads

Download data is not yet available.

Article Details

How to Cite
Dong, K. (2026). Optimizing Cross-Modal Image–Text Matching Algorithms Based on Graph Neural Networks: Implications for Textile Applications. Advanced Electromagnetics, 15(3), 4370–4379. https://doi.org/10.7716/aem.v15i3.3511
Section
Research Articles

References

Y. Cheng, X. Zhu, J. Qian, F. Wen, and P. Liu, “Cross-modal graph matching network for image-text retrieval,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 18, no. 4, pp. 1-23, 2022, doi: 10.1145/3499027.

View Article

S. Qian, D. Xue, Q. Fang, and C. Xu, “Integrating multi-label contrastive learning with dual adversarial graph neural networks for cross-modal retrieval,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 4, pp. 4794-4811, 2022, doi: 10.1109/TPAMI.2022.3188547.

View Article

M. Hou, Z. Zhang, C. Liu, and G. Lu, “Semantic alignment network for multi-modal emotion recognition,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 9, pp. 5318-5329, 2023, doi: 10.1109/TCSVT.2023.3247822.

View Article

R. Raman, P. Das, R. Aggarwal, R. Buch, B. Palanisamy, and T. Basant, “Circular Economy Transitions in Textile, Apparel, and Fashion: AI-Based Topic Modeling and Sustainable Development Goals Mapping,” Sustainability, vol. 17, no. 12, pp. 5342, 2025, doi: 10.3390/su17125342.

View Article

N. Kaur, S. Pandey, and N. Kalra, “Fashion cloth image categorization and retrieval with enhanced intensity using SURF and CNN approach,” International Journal of Clothing Science and Technology, vol. 37, no. 2, pp. 222-241, 2025, doi: 10.1108/IJCST-03-2024-0074.

View Article

T. Yao, Y. Li, Y. Li, G. Wang, and J. Yue, “Cross-modal semantically augmented network for image-text matching,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 20, no. 4, pp. 1-18, 2023, doi: 10.1145/3631356.

View Article

X. Dong, L. Liu, L. Zhu, L. Nie, and H. Zhang, “Adversarial graph convolutional network for cross-modal retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 3, pp. 1634-1645, 2021, doi: 10.1109/TCSVT.2021.3075242.

View Article

K. Mounika and B. V. R. Yadav, “ATTENTION-GUIDED MULTIMODAL GRAPH NETWORKS FOR CROSS-MODAL IMAGE-TEXT MATCHING,” Cuestiones de Fisioterapia, vol. 54, no. 5, pp. 358-372, 2025, doi: 10.48047/69h5ek37.

View Article

Z. Cui, Y. Hu, Y. Sun, J. Gao, and B. Yin, “Cross-modal alignment with graph reasoning for image-text retrieval,” Multimedia Tools and Applications, vol. 81, no. 17, pp. 23615-23632, 2022, doi: 10.1007/s11042-022-12444-8.

View Article

X. Liu, Y. He, M. Cheung Y, X. Xu, and N. Wang, “Learning relationship-enhanced semantic graph for fine-grained image–text matching,” IEEE transactions on cybernetics, vol. 54, no. 2, pp. 948-961, 2022, doi: 10.1109/TCYB.2022.3179020.

View Article

J. Pei, K. Zhong, Z. Yu, L. Wang, and K. Lakshmanna, “Scene graph semantic inference for image and text matching,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 22, no. 5, pp. 1-23, 2023, doi: 10.1145/3563390.

View Article

J. Cao, X. Qin, S. Zhao, and J. Shen, “Bilateral cross-modality graph matching attention for feature fusion in visual question answering,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 3, pp. 4160-4171, 2022, doi: 10.1109/TNNLS.2021.3135655.

View Article

P. Park, S. Jang, Y. Cho, and Y. Kim, “SAM: Cross-modal semantic alignments module for image-text retrieval,” Multimedia Tools and Applications, vol. 83, no. 4, pp. 12363-12377, 2024, doi: 10.1007/s11042-023-15798-9.

View Article

Y. Tao, L. Ma, J. Yu, and H. Zhang, “Memory-based cross-modal semantic alignment network for radiology report generation,” IEEE Journal of Biomedical and Health Informatics, vol. 28, no. 7, pp. 4145-4156, 2024, doi: 10.1109/JBHI.2024.3393018.

View Article

Y. Sun, M. Wang, and Y. Ma, “Semantic-alignment transformer and adversary hashing for cross-modal retrieval,” Applied Intelligence, vol. 54, no. 17, pp. 7581-7602, 2024, doi: 10.1007/s10489-024-05501-2.

View Article

N. Messina, G. Amato, A. Esuli, F. Falchi, C. Gennaro, and S. Marchand-Maillet, “Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 17, no. 4, pp. 1-23, 2021, doi: 10.1145/3451390.

View Article

Y. Hu, K. Wang, M. Liu, H. Tang, and L. Nie, “Semantic collaborative learning for cross-modal moment localization,” ACM Transactions on Information Systems, vol. 42, no. 2, pp. 1-26, 2023, doi: 10.1145/3620669.

View Article

J. Peng S, Y. He, X. Liu, Y. Cheung, X. Xu, and Z. Cui, “Relation-aggregated cross-graph correlation learning for finegrained image–text retrieval,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 2, pp. 2194-2207, 2022, doi: 10.1109/TNNLS.2022.3188569.

View Article

D. Shi, L. Zhu, J. Li, G. Dong, and H. Zhang, “Incomplete cross-modal retrieval with deep correlation transfer,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 20, no. 5, pp. 1-21, 2024, doi: 10.1145/3637442.

View Article

Y. Zeng, Y. Wang, D. Liao, G. Li, W. Huang, J. Xu, et al., “Keyword-based diverse image retrieval with variational multiple instance graph,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 12, pp. 10528-10537, 2022, doi: 10.1109/TNNLS.2022.3168431.

View Article

Z. Li, H. Lu, H. Fu, F. Meng, and G. Gu, “Csan: Cross-coupled semantic adversarial network for cross-modal retrieval,” Artificial Intelligence Review, vol. 58, no. 5, pp. 1-27, 2025, doi: 10.1007/s10462-025-11152-7.

View Article

X. Liang, E. Yang, C. Deng, F. Meng, and G. Gu, “CrossFormer: Cross-modal Representation Learning via Heterogeneous Graph Transformer,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 20, no. 12, pp. 1-21, 2024, doi: 10.1145/3688801.

View Article

Z. Liu, G. Hao, F. Li, X. He, and Y. Zhang, “Multi-modal Sentiment Classification Based On Graph Neural Network And Multi-head Cross-attention Mechanism For Education Emotion Analysis,” Journal of Applied Science and Engineering, vol. 28, no. 6, pp. 1185-1193, 2025, doi: 10.6180/jase.202506_28(6).0002.

View Article

O. Aseniserare and S. Burnett, “Exploring Challenges and Adoption Factors of Business Intelligence Systems in Nigerian Textile and Apparel Industry,” African Journal of Management and Business Research, vol. 19, no. 1, pp. 456-488, 2025, doi: 10.62154/ajmbr.2025.019.01033.

View Article

C. Bermeo-Giraldo M, A. Valencia-Arias, M. Rojas E, and S. Cardona-Acevedo, “Research agenda on the evolution of digital transformation in the textile sector: a bibliometric analysis and research trends,” Discover Sustainability, vol. 6, no. 1, pp. 294, 2025, doi: 10.1007/s43621-025-01091-2.

View Article

L. Weible, T. Domina, R. Dankwa, and L. Agnew, “Recommended base layer ensemble for testing winter outdoor apparel using an infant thermal manikin,” International Journal of Fashion Design, Technology and Education, pp. 1-11, 2025, doi: 10.1080/17543266.2025.2476543.

View Article

J. Guo and J. Ding, “Robust cross-modal retrieval with alignment refurbishment,” Frontiers of Information Technology & Electronic Engineering, vol. 24, no. 10, pp. 1403-1415, 2023, doi: 10.1631/FITEE.2200514.

View Article

N. Han, J. Chen, H. Zhang, H. Wang, and H. Chen, “Adversarial multi-grained embedding network for cross-modal text-video retrieval,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 18, no. 2, pp. 1-23, 2022, doi: 10.1145/3483381.

View Article

Y. Cao, “English text classification model based on graph neural networks and contrastive learning,” International Journal of Information and Communication Technology, vol. 26, no. 25, pp. 48-69, 2025, doi: 10.1504/IJICT.2025.147141.

View Article

Y. Gong, G. Cosma, and A. Finke, “VITR: Augmenting vision transformers with relation-focused learning for cross-modal information retrieval,” ACM Transactions on Knowledge Discovery from Data, vol. 18, no. 9, pp. 1-21, 2024, doi: 10.1145/3686805.

View Article

X. Qin, L. Li, and G. Pang, “Multi-scale motivated neural network for image-text matching,” Multimedia Tools and Applications, vol. 83, no. 2, pp. -4407, 2024, doi: 10.1007/s11042-023-15321-0.

View Article

A. Shin, M. Ishii, and T. Narihira, “Perspectives and prospects on transformer architecture for cross-modal tasks with language and vision,” International journal of computer vision, vol. 130, no. 2, pp. 435-454, 2022, doi: 10.1007/s11263-021-01547-8.

View Article

J. Wang, J. Ke, H. H. Shuai, Y. H. Li, and W. H. Cheng, “Referring expression comprehension via enhanced cross-modal graph attention networks,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 19, no. 2, pp. 1-21, 2023, doi: 10.1145/3548688.

View Article

X. Dong, X. Zhan, Y. Wei, X. Wei, Y. Wang, M. Lu, et al., “Entity-graph enhanced cross-modal pretraining for instance-level product retrieval,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 11, pp. 13117-13133, 2023, doi: 10.1109/TPAMI.2023.3291237.

View Article

S. Xiong, L. Pan, X. Ma, Q. Hu, and E. Beckman, “Unsupervised deep hashing with multiple similarity preservation for cross-modal image-text retrieval,” International Journal of Machine Learning and Cybernetics, vol. 15, no. 10, pp. -4434, 2024, doi: 10.1007/s13042-024-02154-y.

View Article

Y. Chen, C. Hu, T. Kimura, Q. Li, S. Liu, F. Wu, et al., “Semicmt: Contrastive cross-modal knowledge transfer for iot sensing with semi-paired multi-modal signals,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 8, no. 4, pp. 1-30, 2024, doi: 10.1145/3699779.

View Article

H. Reddy A, V. Tirunagari, and V. Rajesh K K, “Analytical Prospective on Fashion Industry: Data-Driven Strategies[M]//Predictive Analytics and Generative AI for Data-Driven Marketing Strategies,” Chapman and Hall/CRC, pp. 33-52, 2024.

N. Rostamzadeh, S. Hosseini, T. Boquet, W. Stokowiec, Y. Zhang, C. Jauvin, et al., “Fashion-gen: The generative fashion dataset and challenge,” arXiv preprint, pp. Jun 21, 2018, doi: 10.48550/arXiv.1806.08317.

View Article

Z. Liu, S. Yan, P. Luo, X. Wang, and X. Tang, “Fashion landmark detection in the wild,” in Leibe, B., Matas, J., Sebe, N., Welling, M. In European Conference on Computer Vision 2016. Cham, Germany: Springer International Publishing. 2016, pp. 229-245, doi: 10.1007/978-3-319-46475-6_15.

View Article

B. Pan, J. Xiang, N. Zhang, and R. Pan, “Two-stage attribute-guided dual attention network for fine-grained fashion retrieval,” Computer Vision and Image Understanding, Art. no. 104497, 2025, doi: 10.1016/j.cviu.2025.104497.

View Article

S. Bell and K. Bala, “Learning visual similarity for product design with convolutional neural networks,” ACM transactions on graphics (TOG), vol. 34, no. 4, pp. 1-10, 2015, doi: 10.1145/2766959.

View Article

W. Yu, X. He, J. Pei, X. Chen, L. Xiong, J. Liu, et al., “Visually aware recommendation with aesthetic features,” The VLDB Journal, vol. 30, no. 4, pp. 495-513, 2021, doi: 10.1007/s00778-021-00651-y.

View Article

Z. Al-Halah, R. Stiefelhagen, and K. Grauman, “Fashion forward: Forecasting visual style in fashion,” In Proceedings of the IEEE international conference on computer vision. Venice, Italy: IEEE, pp.. p. 388-397, 2017, doi: 10.1109/ICCV.2017.50.

View Article

A. Kovashka, D. Parikh, and K. Grauman, “Whittlesearch: Image search with relative attribute feedback,” In 2012 IEEE Conference on Computer Vision and Pattern Recognition. Providence, RI, USA: IEEE, pp.. p. 2973-2980, 2012, doi: 10.1109/CVPR.2012.6248026.

View Article

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > >> 

You may also start an advanced similarity search for this article.