Learning explanations that are hard to vary
Giambattista Parascandolo, Alexander Neitz, Antonio Orvieto, Luigi Gresele, Bernhard Schölkopf
Abstract
In this paper, we investigate the principle that good explanations are hard to vary in the context of deep learning. We show that averaging gradients across examples -akin to a logical OR (_) of patterns -can favor memorization and 'patchwork' solutions that sew together different strategies, instead of identifying invariances. To inspect this, we first formalize a notion of consistency for minima of the loss surface, which measures to what extent a minimum appears only when examples are pooled. We then propose and experimentally validate a simple alternative algorithm based on a logical AND (^), that focuses on invariances and prevents memorization in a set of real-world tasks. Finally, using a synthetic dataset with a clear distinction between invariant and spurious mechanisms, we dissect learning signals and compare this approach to well-established regularizers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6ab8447-3aec-47de-91ff-64aa33a7fc31Cited by top-tier papers65
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville et al.NeurIPS 2021 · 378 citations
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution GeneralizationKartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet et al.NeurIPS 2021 · 372 citations
- Fishr: Invariant Gradient Variances for Out-of-Distribution GeneralizationAlexandre Ramé, Corentin Dancette, Matthieu CordICML 2022 · 262 citations
- Correct-N-Contrast: a Contrastive Approach for Improving Robustness to Spurious CorrelationsMichael Zhang, Nimit Sharad Sohoni, Hongyang R. Zhang, Chelsea Finn et al.ICML 2022 · 230 citations
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 186 citations
Builds on1
Related papers
- The Pitfalls of Memorization: When Memorization Hurts GeneralizationReza Bayat, Mohammad Pezeshki, Elvis Dohmatob, David Lopez-Paz et al.ICLR 2025 · 1 citation
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based OptimizationSatrajit ChatterjeeICLR 2020 · 60 citations
- New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and SoundArushi Gupta, Nikunj Saunshi, Dingli Yu, Kaifeng Lyu et al.NeurIPS 2022 · 12 citations
- Building Reliable Explanations of Unreliable Neural Networks: Locally Smoothing Perspective of Model InterpretationDohun Lim, Hyeonseok Lee, Sungchan KimCVPR 2021
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
