ICML2026

Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks

Xueting Chen, Jun-Jie Huang, Yan Yan, Long Lan, Yuhua Tang, Wenjing Yang

摘要

Multi-modal prompt learning is a parameter-efficient approach to adapting large vision-language models to downstream classification tasks. However, prompts can inadvertently evolve into a high-capacity pathway that encodes environment-dependent spurious correlations, which are predictive only in the source domain and thereby undermine transferability. To address this issue, this paper introduces Do-Prompt , a compress-and-intervene framework that brings together variational bottlenecks and causal interventions for robust prompt tuning. We model prompts as stochastic latent variables and impose a variational prompt bottleneck to explicitly regulate the information transmitted through prompts, effectively mitigating their tendency to memorize spurious nuisance cues. Building on this capacity constraint, we propose lightweight prompt-level interventions by perturbing the environment-related prompt components and enforcing prediction consistency under these do -style perturbations. This synergistic integration encourages reliance on task-stable, invariant semantics rather than spurious prompt content. Notably, Do-Prompt is plug-and-play compatible with existing multi-modal prompt tuning pipelines and introduces negligible computational overhead. Extensive experiments on base-to-novel generalization, cross-dataset transfer, and ImageNet distribution shifts demonstrate consistent performance gains, with particularly notable improvements on datasets exhibiting pronounced domain or texture biases.