Convolutional Visual Prompt for Robust Visual Perception
Yun-Yun Tsai, Chengzhi Mao, Junfeng Yang
Abstract
Vision models are often vulnerable to out-of-distribution (OOD) samples, which often need adaptation to fix them. While visual prompts offer a lightweight method of input-space adaptation for large-scale vision models, they rely on a high-dimensional additive vector and labeled data. This leads to overfitting when adapting models in a self-supervised test-time setting without labels. We introduce convolutional visual prompts (CVP) for label-free test-time adaptation for robust visual perception. The structured nature of CVP demands fewer trainable parameters, less than 1% compared to standard visual prompts, combating overfitting. Extensive experiments and analysis on a wide variety of OOD visual perception tasks show that our approach is effective, improving robustness by up to 5.87% over several large-scale models. * equal contributions 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a51d28c7-3d34-4b3a-bae7-cd561aff6506Cited by top-tier papers13
- AutoVP: An Automated Visual Prompting Framework and BenchmarkHsi-Ai Tsao, Lei Hsiung, Pin-Yu Chen, Si Liu et al.ICLR 2024 · 29 citations
- Convolutional Prompting meets Language Models for Continual LearningAnurag Roy, Riddhiman Moulick, Vinay Kumar Verma, Saptarshi Ghosh et al.CVPR 2024 · 15 citations
- Black-Box Test-Time Prompt Tuning for Vision-Language ModelsFan'an Meng, Chaoran Cui, Hongjun Dai, Shuai GongAAAI 2025 · 6 citations
- Correlated Low-Rank Adaptation for ConvNetsWu Ran, Weijia Zhang, Shuyang Pang, Qi Zhu et al.NeurIPS 2025 · 5 citations
- Efficient Deformable Convolutional Prompt for Continual Test-Time Adaptation in Medical Image SegmentationShiyu Liu, Daoqiang Zhang, Xiaoke HaoAAAI 2025 · 4 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- VPA: Fully Test-Time Visual Prompt AdaptationJiachen Sun, Mark Ibrahim, Melissa Hall, Ivan Evtimov et al.ACM MM 2023 · 7 citations
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting MitigationShohei EnomotoNeurIPS 2025 · 1 citation
- Exploring Sparse Visual Prompt for Domain Adaptive Dense PredictionSenqiao Yang, Jiarui Wu, Jiaming Liu, Xiaoqi Li et al.AAAI 2024 · 38 citations
- Decorate the Newcomers: Visual Domain Prompt for Continual Test Time AdaptationYulu Gan, Yan Bai, Yihang Lou, Xianzheng Ma et al.AAAI 2023 · 145 citations
- Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIP's Zero-Shot GeneralizationRan Wang, Hua Zuo, Zhen Fang, Jie LuACM MM 2024 · 7 citations
