Fishr: Invariant Gradient Variances for Out-of-Distribution Generalization
Alexandre Ramé, Corentin Dancette, Matthieu Cord
Abstract
Learning robust models that generalize well under changes in the data distribution is critical for real-world applications. To this end, there has been a growing surge of interest to learn simultaneously from multiple training domains - while enforcing different types of invariance across those domains. Yet, all existing approaches fail to show systematic benefits under controlled evaluation protocols. In this paper, we introduce a new regularization - named Fishr - that enforces domain invariance in the space of the gradients of the loss: specifically, the domain-level variances of gradients are matched across training domains. Our approach is based on the close relations between the gradient covariance, the Fisher Information and the Hessian of the loss: in particular, we show that Fishr eventually aligns the domain-level loss landscapes locally around the final weights. Extensive experiments demonstrate the effectiveness of Fishr for out-of-distribution generalization. Notably, Fishr improves the state of the art on the DomainBed benchmark and performs consistently better than Empirical Risk Minimization. Our code is available at https://github.com/alexrame/fishr.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4585064d-7973-44ce-9176-c91b3e9bedcaCited by top-tier papers101
- Prompt-aligned Gradient for Prompt TuningBeier Zhu, Yulei Niu, Yucheng Han, Yue Wu et al.ICCV 2023 · 475 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- Ensemble of Averages: Improving Model Selection and Boosting Performance in Domain GeneralizationDevansh Arpit, Huan Wang, Yingbo Zhou, Caiming XiongNeurIPS 2022 · 232 citations
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy et al.NeurIPS 2022 · 183 citations
- Local Learning Matters: Rethinking Data Heterogeneity in Federated LearningMatías Mendieta, Taojiannan Yang, Pu Wang, Minwoo Lee et al.CVPR 2022 · 176 citations
Builds on28
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
Related papers
- Gradient Matching for Domain GeneralizationYuge Shi, Jeffrey Seely, Philip H. S. Torr, Siddharth Narayanaswamy et al.ICLR 2022 · 358 citations
- Understanding Hessian Alignment for Domain GeneralizationSobhan Hemati, Guojun Zhang, Amir Hossein Estiri, Xi ChenICCV 2023 · 21 citations
- Adaptive Risk Minimization: Learning to Adapt to Domain ShiftMarvin Zhang, Henrik Marklund, Nikita Dhawan, Abhishek Gupta et al.NeurIPS 2021 · 284 citations
- Lost Domain Generalization Is a Natural Consequence of Lack of Training DomainsYimu Wang, Yihan Wu, Hongyang ZhangAAAI 2024 · 7 citations
- An Empirical Investigation of Domain Generalization with Empirical Risk MinimizersRamakrishna Vedantam, David Lopez-Paz, David J. SchwabNeurIPS 2021 · 49 citations
