Discrete Latent Features Ablate Adversarial Attack: A Robust Prompt Tuning Framework for VLMs
Yang Chen, Yanbin Wei, James T. Kwok, Yu Zhang
Abstract
While adversarial fine-tuning can enhance the robustness of vision-language models (VLMs), such approaches are computationally expensive. Adversarial prompt tuning has emerged as a practical alternative. However, existing methods are limited by their reliance on vulnerable continuous image features. To mitigate the vulnerability in the feature representation, we propose DEFEAT (Discrete LatEnt FeaturE based Adversarial Training), a robust prompt tuning framework for VLMs. Specifically, the DEFEAT method introduces a perturbation discrete shield module that reconstructs discrete latent features and designs a logits fusion strategy, substantially reducing the discrepancy between clean and adversarial image representations. Moreover, the DEFEAT method integrates prompt tuning with adversarial training while applying regularization from learnable prompts to hand-crafted prompts, further enhancing the adversarial robustness. Extensive experiments across 15 datasets validate the effectiveness of the proposed DEFEAT method among existing adversarial prompt tuning methods. The official code is available at https://github.com/cheny02/DEFEAT-ICLR2026 . * Equal contribution. † Corresponding author. RELATED WORK CLIP-based VLMs. VLMs have significantly boosted cognitive capabilities by merging visual and textual modalities, excelling in real-world vision tasks (Liu et al., 2023; Zhu et al., 2024) . The introduction of CLIP (Radford et al., 2021) , trained on about 400 million image-text pairs, was particularly transformative, establishing a new paradigm for vision-language representation learning. Numerous subsequent works have followed this paradigm, proposing a broad family of CLIP-like models, including ALIGN (Jia et al., 2021), EVA-CLIP (Sun et al., 2023), OpenCLIP (Ilharco et al.,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e392bee6-5c56-470c-a4e2-f67a2656deb1Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Learning Robust Vision-Language Models from Natural Latent SpacesZhangyun Wang, Ni Ding, Aniket MahantiNeurIPS 2025 · 3 citations
- One Prompt Word is Enough to Boost Adversarial Robustness for Pre-Trained Vision-Language ModelsLin Li, Haoyan Guan, Jianing Qiu, Michael W. SpratlingCVPR 2024
- R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt TuningLijun Sheng, Jian Liang, Zilei Wang, Ran HeCVPR 2025
- TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsXin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen et al.CVPR 2025
- Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language ModelsWonjeong Choi, Sejong Ryu, Jungmoon Lee, Dong-Jun Han et al.ICLR 2026
