Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
Yehonatan Elisha, Oren Barkan, Noam Koenigstein
摘要
Vision Transformers (ViTs) often degrade under distribution shifts because they rely on spurious correlations, such as background cues, rather than semantically meaningful features. Existing regularization methods, typically relying on simple foreground-background masks, which fail to capture the fine-grained semantic concepts that define an object (e.g., "long beak" and "wings" for a "bird"). As a result, these methods provide limited robustness to distribution shifts. To address this limitation, we introduce a novel finetuning framework that steers model reasoning toward concept-level semantics. Our approach optimizes the model's internal relevance maps to align with spatially grounded concept masks. These masks are generated automatically, without manual annotation: class-relevant concepts are first proposed using an LLM-based, label-free method, and then segmented using a VLM. The finetuning objective aligns relevance with these concept regions while simultaneously suppressing focus on spurious background areas. Notably, this process requires only a minimal set of images and uses half of the dataset classes. Extensive experiments on five out-of-distribution benchmarks demonstrate that our method improves robustness across multiple ViT-based models. Furthermore, we show that the resulting relevance maps exhibit stronger alignment with semantic object parts, offering a scalable path toward more robust and interpretable vision models. Finally, we confirm that concept-guided masks provide more effective supervision for model robustness than conventional segmentation maps, supporting our central hypothesis. Our code is provided at:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- Optimizing Relevance Maps of Vision Transformers Improves RobustnessHila Chefer, Idan Schwartz, Lior WolfNeurIPS 2022 · 被引用 55 次
- ERICT: Enhancing Robustness by Identifying Concept Tokens in Zero-Shot Vision Language ModelsXinpeng Dong, Min Zhang, Didi Zhu, Ye Jun Jian 等ICML 2025
- Locality Alignment Improves Vision-Language ModelsIan Connick Covert, Tony Sun, James Zou, Tatsunori HashimotoICLR 2025
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision TransformersLi Ren, Chen Chen, Liqiang Wang, Kien A. HuaCVPR 2025
- MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationZhiwei Yang, Yucong Meng, Kexue Fu, Shuo Wang 等AAAI 2025 · 被引用 14 次
