LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual Images
Viraj Prabhu, Sriram Yenamandra, Prithvijit Chattopadhyay, Judy Hoffman
Abstract
We propose an automated algorithm to stress-test a trained visual model by generating language-guided counterfactual test images (LANCE). Our method leverages recent progress in large language modeling and text-based image editing to augment an IID test set with a suite of diverse, realistic, and challenging test images without altering model weights. We benchmark the performance of a diverse set of pre-trained models on our generated data and observe significant and consistent performance drops. We further analyze model sensitivity across different types of edits, and demonstrate its applicability at surfacing previously unknown class-level model biases in ImageNet. Code is available at https://github.com/virajprabhu/lance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0932f81-488c-4568-a836-fac528a21e76Cited by top-tier papers10
- A Sober Look at the Robustness of CLIPs to Spurious FeaturesQizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt et al.NeurIPS 2024 · 46 citations
- Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingYinuo Jing, Ruxu Zhang, Kongming Liang, Yongxiang Li et al.NeurIPS 2024 · 13 citations
- Visual Data Diagnosis and Debiasing with Concept GraphsRwiddhi Chakraborty, Yinong Wang, Jialu Gao, Runkai Zheng et al.NeurIPS 2024 · 9 citations
- Domain Gap Embeddings for Generative Dataset AugmentationYinong Oliver Wang, Younjoon Chung, Chen Henry Wu, Fernando De la TorreCVPR 2024 · 8 citations
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng et al.NeurIPS 2025 · 7 citations
Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
Related papers
- Can We Debias Multimodal Large Language Models via Model Editing?Zecheng Wang, Xinye Li, Zhanyue Qin, Chunshan Li et al.ACM MM 2024 · 2 citations
- Diversify Your Vision Datasets with Automatic Diffusion-based AugmentationLisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang et al.NeurIPS 2023 · 136 citations
- DISCO: Distilling Counterfactuals with Large Language ModelsZeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal et al.ACL 2023 · 27 citations
- Dually Self-Improved Counterfactual Data Augmentation Using Large Language ModelLuhao Zhang, Xinyu Zhang, Linmei Hu, Dandan Song et al.ACL 2025 · 1 citation
- Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image ClassificationZequn Zeng, Yudi Su, Jianqiao Sun, Tiansheng Wen et al.CVPR 2025
