Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization
Elan Rosenfeld, Andrej Risteski
摘要
We identify a new phenomenon in neural network optimization which arises from the interaction of depth and a particular heavy-tailed structure in natural data. Our result offers intuitive explanations for several previously reported observations about network training dynamics. In particular, it implies a conceptually new cause for progressive sharpening and the edge of stability; we also highlight connections to other concepts in optimization and generalization including grokking, simplicity bias, and Sharpness-Aware Minimization. Experimentally, we demonstrate the significant influence of paired groups of outliers in the training data with strong opposing signals: consistent, large magnitude features which dominate the network output throughout training and provide gradients which point in opposite directions. Due to these outliers, early optimization enters a narrow valley which carefully balances the opposing groups; subsequent sharpening causes their loss to rise rapidly, oscillating between high on one group and then the other, until the overall loss spikes. We describe how to identify these groups, explore what sets them apart, and carefully study their effect on the network's optimization and behavior. We complement these experiments with a mechanistic explanation on a toy example of opposing signals and a theoretical analysis of a two-layer linear network on a simple model. Our finding enables new qualitative predictions of training behavior which we confirm experimentally. It also provides a new lens through which to study and improve modern training practices for stochastic optimization, which we highlight via a case study of Adam versus SGD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Iteration Head: A Mechanistic Study of Chain-of-ThoughtVivien Cabannes, Charles Arnal, Wassim Bouaziz, Xingyu Yang 等NeurIPS 2024 · 被引用 44 次
- Decomposing and Editing Predictions by Modeling Model ComputationHarshay Shah, Andrew Ilyas, Aleksander MadryICML 2024 · 被引用 25 次
- Hidden Breakthroughs in Language Model TrainingSara Kangaslahti, Elan Rosenfeld, Naomi SaphraICLR 2026 · 被引用 17 次
- Learning Associative Memories with Gradient DescentVivien Cabannes, Berfin Simsek, Alberto BiettiICML 2024 · 被引用 13 次
- Deep sequence models tend to memorize geometrically; it is unclear whyShahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv KumarICML 2026 · 被引用 11 次
它引用的顶会 Paper24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
相关 Paper
- Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and MoreGeonhui Yoo, Minhak Song, Chulhee YunICML 2025
- A Minimalist Example of Edge-of-Stability and Progressive SharpeningLiming Liu, Zixuan Zhang, Simon S. Du, Tuo ZhaoNeurIPS 2025 · 被引用 4 次
- Egalitarian Gradient Descent: A Simple Approach to Accelerated GrokkingAli Saheb Pasand, Elvis DohmatobICLR 2026 · 被引用 1 次
- Omnigrok: Grokking Beyond Algorithmic DataZiming Liu, Eric J. Michaud, Max TegmarkICLR 2023 · 被引用 8 次
- Grokking at the Edge of Linear SeparabilityAlon Beck, Noam Itzhak Levi, Yohai Bar-SinaiICML 2025
