Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks
Xueting Chen, Jun-Jie Huang, Yan Yan, Long Lan, Yuhua Tang, Wenjing Yang
Abstract
Multi-modal prompt learning is a parameter-efficient approach to adapting large vision-language models to downstream classification tasks. However, prompts can inadvertently evolve into a high-capacity pathway that encodes environment-dependent spurious correlations, which are predictive only in the source domain and thereby undermine transferability. To address this issue, this paper introduces Do-Prompt , a compress-and-intervene framework that brings together variational bottlenecks and causal interventions for robust prompt tuning. We model prompts as stochastic latent variables and impose a variational prompt bottleneck to explicitly regulate the information transmitted through prompts, effectively mitigating their tendency to memorize spurious nuisance cues. Building on this capacity constraint, we propose lightweight prompt-level interventions by perturbing the environment-related prompt components and enforcing prediction consistency under these do -style perturbations. This synergistic integration encourages reliance on task-stable, invariant semantics rather than spurious prompt content. Notably, Do-Prompt is plug-and-play compatible with existing multi-modal prompt tuning pipelines and introduces negligible computational overhead. Extensive experiments on base-to-novel generalization, cross-dataset transfer, and ImageNet distribution shifts demonstrate consistent performance gains, with particularly notable improvements on datasets exhibiting pronounced domain or texture biases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cddd0755-6d44-46fd-9189-2f140e86d283Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
Related papers
- VPA: Fully Test-Time Visual Prompt AdaptationJiachen Sun, Mark Ibrahim, Melissa Hall, Ivan Evtimov et al.ACM MM 2023 · 7 citations
- Prompt Tuning for CLIP on the Pretrained ManifoldXi Yang, Yuanrong Xu, Weigang Zhang, Guangming Lu et al.ICML 2026 · 1 citation
- Domain-Agnostic Mutual Prompting for Unsupervised Domain AdaptationZhekai Du, Xinyao Li, Fengling Li, Ke Lu et al.CVPR 2024
- Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language ModelsJie Zhang, Xiaosong Ma, Song Guo, Peng Li et al.ICML 2024 · 10 citations
- LaViP: Language-Grounded Visual PromptingNilakshan Kunananthaseelan, Jing Zhang, Mehrtash HarandiAAAI 2024 · 6 citations
