Understanding and Improving Feature Learning for Out-of-Distribution Generalization
Yongqiang Chen, Wei Huang, Kaiwen Zhou, Yatao Bian, Bo Han, James Cheng
摘要
A common explanation for the failure of out-of-distribution (OOD) generalization is that the model trained with empirical risk minimization (ERM) learns spurious features instead of invariant features. However, several recent studies challenged this explanation and found that deep networks may have already learned sufficiently good features for OOD generalization. Despite the contradictions at first glance, we theoretically show that ERM essentially learns both spurious and invariant features, while ERM tends to learn spurious features faster if the spurious correlation is stronger. Moreover, when fed the ERM learned features to the OOD objectives, the invariant feature learning quality significantly affects the final OOD performance, as OOD objectives rarely learn new features. Therefore, ERM feature learning can be a bottleneck to OOD generalization. To alleviate the reliance, we propose Feature Augmented Training (FeAT), to enforce the model to learn richer features ready for OOD generalization. FeAT iteratively augments the model to learn new features while retaining the already learned features. In each round, the retention and augmentation operations are performed on different subsets of the training data that capture distinct features. Extensive experiments show that FeAT effectively learns richer features thus boosting the performance of various OOD objectives 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo 等NeurIPS 2024 · 被引用 74 次
- A Sober Look at the Robustness of CLIPs to Spurious FeaturesQizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt 等NeurIPS 2024 · 被引用 46 次
- Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and MethodBikang Pan, Wei Huang, Ye ShiNeurIPS 2024 · 被引用 28 次
- FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model FusionZhenheng Tang, Yonggang Zhang, Peijie Dong, Yiu-ming Cheung 等NeurIPS 2024 · 被引用 28 次
- On the Comparison between Multi-modal and Single-modal Contrastive LearningWei Huang, Andi Han, Yongqiang Chen, Yuan Cao 等NeurIPS 2024 · 被引用 26 次
它引用的顶会 Paper44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
相关 Paper
- On the Connection between Invariant Learning and Adversarial Training for Out-of-Distribution GeneralizationShiji Xin, Yifei Wang, Jingtong Su, Yisen WangAAAI 2023 · 被引用 14 次
- Simple and Fast Group Robustness by Automatic Feature ReweightingShikai Qiu, Andres Potapczynski, Pavel Izmailov, Andrew Gordon WilsonICML 2023 · 被引用 78 次
- Sparse Invariant Risk MinimizationXiao Zhou, Yong Lin, Weizhong Zhang, Tong ZhangICML 2022 · 被引用 85 次
- Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to GeneralizeTianren Zhang, Chujie Zhao, Guanyu Chen, Yizhou Jiang 等ICML 2024 · 被引用 12 次
- DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic AugmentationHaoyue Bai, Rui Sun, Lanqing Hong, Fengwei Zhou 等AAAI 2021 · 被引用 88 次
