On Guiding Visual Attention with Language Specification
Suzanne Petryk, Lisa Dunlap, Keyan Nasseri, Joseph Gonzalez, Trevor Darrell, Anna Rohrbach
Abstract
While real world challenges typically define visual categories with language words or phrases, most visual classification methods define categories with numerical indices. However, the language specification of the classes provides an especially useful prior for biased and noisy datasets, where it can help disambiguate what features are task-relevant. Recently, large-scale multimodal models have been shown to recognize a wide variety of high-level concepts from a language specification even without additional image training data, but they are often unable to distinguish classes for more fine-grained tasks. CNNs, in contrast, can extract subtle image features that are required for fine-grained discrimination, but will overfit to any bias or noise in datasets. Our insight is to use high-level language specification as advice for constraining the classification evidence to task-relevant features, instead of distractors. To do this, we ground task-relevant words or phrases with attention maps from a pretrained large-scale model. We then use this grounding to supervise a classifier's spatial attention away from distracting context. We show that supervising spatial attention in this way improves performance on classification tasks with biased and noisy data, including 3 −15% worst-group accuracy improvements and 41-45% relative improvements on fairness metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fa28a01-3b8a-4eb6-b308-9a14a2335633Cited by top-tier papers13
- Diversify Your Vision Datasets with Automatic Diffusion-based AugmentationLisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang et al.NeurIPS 2023 · 136 citations
- Mitigating Spurious Correlations in Multi-modal Models during Fine-tuningYu Yang, Besmira Nushi, Hamid Palangi, Baharan MirzasoleimanICML 2023 · 65 citations
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 22 citations
- B-cosification: Transforming Deep Neural Networks to be Inherently InterpretableShreyash Arya, Sukrut Rao, Moritz Böhle, Bernt SchieleNeurIPS 2024 · 14 citations
- Calibrating Multi-modal Representations: A Pursuit of Group Robustness without AnnotationsChenyu You, Yifei Min, Weicheng Dai, Jasjeet S. Sekhon et al.CVPR 2024 · 10 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele et al.NeurIPS 2020 · 392 citations
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
Related papers
- SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal BiasWenqian Ye, Di Wang, Guangtao Zheng, Bohan Liu et al.AAAI 2026
- On Large Multimodal Models as Open-World Image ClassifiersAlessandro Conti, Massimiliano Mancini, Enrico Fini, Yiming Wang et al.ICCV 2025 · 3 citations
- Learning Concise and Descriptive Attributes for Visual RecognitionAn Yan, Yu Wang, Yiwu Zhong, Chengyu Dong et al.ICCV 2023 · 94 citations
- See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object UnderstandingBoyuan Sun, Bo-Wen Yin, Yuan-Ming Li, Xihan Wei et al.CVPR 2026 · 1 citation
- Visual Classification via Description from Large Language ModelsSachit Menon, Carl VondrickICLR 2023 · 57 citations
