Masked Images Are Counterfactual Samples for Robust Fine-Tuning
Yao Xiao, Ziyi Tang, Pengxu Wei, Cong Liu, Liang Lin
摘要
Deep learning models are challenged by the distribution shift between the training data and test data. Recently, the large models pre-trained on diverse data have demonstrated unprecedented robustness to various distribution shifts. However, fine-tuning these models can lead to a trade-off between in-distribution (ID) performance and out-of-distribution (OOD) robustness. Existing methods for tackling this trade-off do not explicitly address the OOD robustness problem. In this paper, based on causal analysis of the aforementioned problems, we propose a novel fine-tuning method, which uses masked images as counterfactual samples that help improve the robustness of the fine-tuning model. Specifically, we mask either the semantics-related or semantics-unrelated patches of the images based on class activation map to break the spurious correlation, and refill the masked patches with patches from other images. The resulting counterfactual samples are used in feature-based distillation with the pre-trained model. Extensive experiments verify that regularizing the fine-tuning with the proposed masked images can achieve a better trade-off between ID and OOD performance, surpassing previous methods on the OOD performance. Our code is available at https://github.com/Coxy7/robust-finetuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DGMamba: Domain Generalization via Generalized State Space ModelShaocong Long, Qianyu Zhou, Xiangtai Li, Xuequan Lu 等ACM MM 2024 · 被引用 15 次
- Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained ModelsPeiyan Zhang, Haoyang Liu, Chaozhuo Li, Xing Xie 等ICLR 2024 · 被引用 8 次
- Causal-JEPA: Learning World Models through Object-Level Latent MaskingHeejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun 等ICML 2026 · 被引用 7 次
- Contrastive Learning Relies More on Spatial Inductive Bias Than Supervised Learning: An Empirical StudyYuanyi Zhong, Haoran Tang, Jun-Kun Chen, Yu-Xiong WangICCV 2023 · 被引用 4 次
- Decompose-and-Compose: A Compositional Approach to Mitigating Spurious CorrelationFahimeh Hosseini Noohdani, Parsa Hosseini, Aryan Yazdan Parast, Hamidreza Yaghoubi Araghi 等CVPR 2024 · 被引用 3 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic AugmentationHaoyue Bai, Rui Sun, Lanqing Hong, Fengwei Zhou 等AAAI 2021 · 被引用 88 次
- MASKER: Masked Keyword Regularization for Reliable Text ClassificationSeung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee 等AAAI 2021 · 被引用 39 次
- Teaching Small Language Models Reasoning through Counterfactual DistillationTao Feng, Yicheng Li, Chenglin Li, Hao Chen 等EMNLP 2024 · 被引用 1 次
- Debiased Fine-Tuning for Vision-Language Models by Prompt RegularizationBeier Zhu, Yulei Niu, Saeil Lee, Minhoe Hur 等AAAI 2023 · 被引用 34 次
- Contrastive Out-of-Distribution Detection for Pretrained TransformersWenxuan Zhou, Fangyu Liu, Muhao ChenEMNLP 2021 · 被引用 63 次
