Sparse Adversarial Patch Attack and Robustness Evaluation Algorithm for Vision-Language Models

Main Article Content

T. Y. Chen
X. Y. Hu
J. F. Wang
J. S. Xiao

Abstract

Visual language models (VLMs) have demonstrated outstanding performance in high-value domains such as autonomous driving, unmanned system navigation, and intelligent question-answering; however, the security of their cross-modal alignment mechanisms has not yet been fully verified. Existing visual adversarial patch attacks typically rely on continuous, dense pixel perturbations, which are easily detected and blocked by anomaly detection systems in practical engineering applications. This paper proposes a novel sparse adversarial patch attack algorithm (Sparse Patch Attack, SPA), which successfully misleads the text generation results of VLMs by generating highly dispersed discrete pixel perturbations in non-salient regions of the image. To achieve this, we introduce a differentiable L0-norm approximation and a cross-attention masking mechanism to minimize the number of modified pixels. Furthermore, addressing the characteristics of large-scale model open-ended text generation, we construct a multi-dimensional robustness evaluation framework covering semantic deviation, target achievement rate, and visual concealment. Preliminary experiments on mainstream visual-language models (such as LLaVA and BLIP-2) demonstrate that the SPA algorithm can achieve high success rates in targeted cross-modal attacks with an extremely low pixel modification rate (<1%). This study reveals a novel security vulnerability in visual-language models within complex real-world environments and provides a quantitative evaluation benchmark for future defense mechanisms in multimodal models.

Downloads

Download data is not yet available.

Article Details

How to Cite
Chen, T. Y., Hu, X. Y., Wang, J. F., & Xiao, J. S. (2026). Sparse Adversarial Patch Attack and Robustness Evaluation Algorithm for Vision-Language Models. Advanced Electromagnetics, 15(3), 10779–10785. https://doi.org/10.7716/aem.v15i3.4284
Section
Research Articles

References

H. Liu, C. Li, Q. Wu, and Y. J. Lee, "Visual instruction tuning," in Advances in Neural Information Processing Systems, vol. 36, 2023.

J. Li, D. Li, S. Savarese, and S. Hoi, "BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models," in Int. Conf. Mach. Learn., PMLR, 2023, pp. 19730–19742.

X. Zhou, M. Liu, E. Yurtsever, and Z. L. Zhang, "Vision language models in autonomous driving: A survey and outlook," arXiv preprint arXiv:2310.14414, 2023.

S. Chen, Z. Liu, Q. Zhu, and H. Ma, "Exploring embodied multimodal large models," arXiv preprint arXiv:2502.15336, 2025.

J. C. Costa, T. Roxo, H. Proenca, and P. N. Inacio, "How deep learning sees the world: A survey on adversarial attacks and defenses," arXiv preprint arXiv:2305.10862, 2023.

H. Wei, J. Wang, X. Jia, and Y. Yang, "Physical adversarial attack meets computer vision: A decade survey," IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12, pp. 1–20, 2024.

D. Liu, M. Wang, J. Zhang, and J. Xu, "A survey of attacks on large vision-language models," arXiv preprint arXiv:2407.07403, 2024.

D. Kong, S. Qiu, J. Ma, and Y. Chen, "Patch is enough: Naturalistic adversarial patch against vision-language pre-training models," Vis. Intell., vol. 2, p. 33, 2024.

Q. Guo, X. Jia, S. Pang, and X. Zhang, "PhysPatch: A physically realizable and transferable adversarial patch attack for multimodal large language models-based autonomous driving systems," arXiv preprint arXiv:2508.05167, 2025.

E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, "Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models," in Int. Conf. Learn. Represent., 2024.

S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, "Distributed optimization and statistical learning via the alternating direction method of multipliers," Found. Trends Mach. Learn., vol. 3, no. 1, pp. 1–122, 2011.

A. Athalye, L. Engstrom, A. Ilyas, and K. Kwok, "Synthesizing robust adversarial examples," in Int. Conf. Mach. Learn., PMLR, 2018, pp. 284–293.

K. Kunanbayev, V. Shen, and D. S. Kim, "Training ViT with limited data for Alzheimer’s disease classification: An empirical study," in Med. Image Comput. Comput. Assist. Interv., Cham: Springer, vol. 15012, 2024.

H. Yang, M. Xu, Z. Sun, and J. Chen, "CLIP-ViT detector: Side adapter with prompt for vision transformer object detection," in Comput. Artif. Intell., 2024, pp. 1–8.

S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, "Transformers in vision: A survey," ACM Comput. Surv., vol. 54, no. 10s, pp. 1–41, 2022.

K. Steunou, T. Druilhe, and S. Saue, "Sparse representations improve adversarial robustness of neural network classifiers," arXiv preprint arXiv:2509.21130, 2025.

W. Nam, J. Kim, and D. Kim, "RISOPA: Rapid imperceptible strong one-pixel attacks in deep neural networks," Mathematics, vol. 12, no. 7, p. 1083, 2024.

K. Ding, K. Ma, S. Wang, and D. Fang, "Image quality assessment: Unifying structure and texture similarity," IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 5, pp. 2567–2581, 2022.

H. Liu, C. Li, Y. Li, and Y. J. Lee, "Improved baselines with visual instruction tuning," arXiv preprint arXiv:2310.03744, 2023.

W. Dai, J. Li, D. Li, A. M. H. Tiong, and S. Hoi, "InstructBLIP: Towards general-purpose vision-language models with instruction tuning," in Adv. Neural Inf. Process. Syst., vol. 36, 2023.

D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, "MiniGPT-4: Enhancing vision-language understanding with advanced large language models," arXiv preprint arXiv:2304.10592, 2023.

Z. Yang, Z. Gan, J. Wang, and X. Hu, "Unified vision-language pre-training for image captioning and VQA," Proc. AAAI Conf. Artif. Intell., vol. 36, no. 11, pp. 13041–13049, 2022.

F. Ishmam, A. Sadeghzadeh, R. Krishna, and M. Sadeghzadeh, "Visual robustness benchmark for visual question answering (VQA)," in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis., 2025.

N. Zhang, W. Tao, X. Xiao, and X. Chen, "Attention-guided patch-wise sparse adversarial attacks on vision-language-action models," arXiv preprint arXiv:2511.21663, 2025.

J. S. Xiao, H. Ma, R. D. Chen, and H. J. Pan, "STKPS-Net: Spatio-Temporal Key Patch Selection Network for few-shot anomalous action recognition," IEEE Trans. Inf. Forensics Security, vol. 21, pp. 827–838, 2026.