Prompting is a Double-Edged Sword: Improving Worst-Group Robustness of Foundation Models
Amrith Setlur, Saurabh Garg, Virginia Smith, Sergey Levine
Abstract
Machine learning models fail catastrophically under distribution shift, but a surprisingly effective way to empirically improve robustness to some types of shift (e.g., Imagenet-A/C) is to use stronger open-vocabulary classifiers derived from foundation models. In this work, we first note that for shifts governed by spurious correlations (features spuriously correlated with the label on the training data, but not on test), the zero-shot and few-shot performance of foundation models is no better than ERM models, and remains unchanged when pretrained data/model size is scaled. Secondly, even in these situations, foundation models are quite accurate at predicting the value of the spurious feature. In a simplified setup, we theoretically analyze both these findings. Specifically, we show that during contrastive pretraining, the simplicity bias of foundation models tends to result in the learning of features that mostly rely on the spurious attribute, compared to more robust features. We leverage these observations to propose Prompting for Robustness (PfR) which first uses foundation models to zero-shot predict the spurious attribute on labeled examples, and then learns a classifier with balanced performance across different groups of labels and spurious attribute. Across 5 vision and language tasks, we show that PfR's performance nearly equals that of an oracle algorithm (group DRO) that leverages human labeled spurious attributes 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96f7460f-cdcc-4746-a9d0-692bfbae823aCited by top-tier papers1
Ask how each one uses itBuilds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Examining and Combating Spurious Features under Distribution ShiftChunting Zhou, Xuezhe Ma, Paul Michel, Graham NeubigICML 2021 · 78 citations
- Think Twice: Test-Time Reasoning for Robust CLIP Zero-Shot ClassificationShenyu Lu, Zhaoying Pan, Xiaoqian WangICCV 2025 · 1 citation
- Spurious Correlation-Aware Embedding Regularization for Worst-Group RobustnessSubeen Park, JOOWANG KIM, Hakyung Lee, Sunjae yoo et al.ICLR 2026 · 2 citations
- Class-Conditional Distribution Balancing for Group Robust ClassificationMiaoyun Zhao, Qiang ZhangICML 2026 · 1 citation
- ERICT: Enhancing Robustness by Identifying Concept Tokens in Zero-Shot Vision Language ModelsXinpeng Dong, Min Zhang, Didi Zhu, Ye Jun Jian et al.ICML 2025
