Do We Need Zero Training Loss After Achieving Zero Training Error?
Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu, Masashi Sugiyama
摘要
Overparameterized deep networks have the capacity to memorize training data with zero training error. Even after memorization, the training loss continues to approach zero, making the model overconfident and the test performance degraded. Since existing regularizers do not directly aim to avoid zero training loss, they often fail to maintain a moderate level of training loss, ending up with a too small or too large loss. We propose a direct solution called flooding that intentionally prevents further reduction of the training loss when it reaches a reasonably small value, which we call the flooding level. Our approach makes the loss float around the flooding level by doing mini-batched gradient descent as usual but gradient ascent if the training loss is below the flooding level. This can be implemented with one line of code, and is compatible with any stochastic optimizer and other regularizers. With flooding, the model will continue to "random walk" with the same non-zero training loss, and we expect it to drift into an area with a flat loss landscape that leads to better generalization. We experimentally show that flooding improves performance and as a byproduct, induces a double descent curve of the test loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Understanding and Improving Early Stopping for Learning with Noisy LabelsYingbin Bai, Erkun Yang, Bo Han, Yanhua Yang 等NeurIPS 2021 · 被引用 307 次
- Discriminative Complementary-Label Learning with Weighted LossYi Gao, Min-Ling ZhangICML 2021 · 被引用 48 次
- Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-TuningQin Liu, Rui Zheng, Bao Rong, Jingyi Liu 等ACL 2022 · 被引用 35 次
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and ImplicationsZiqiao Wang, Yongyi MaoICLR 2022 · 被引用 33 次
- Coupled Confusion Correction: Learning from Crowds with Sparse AnnotationsHansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan 等AAAI 2024 · 被引用 23 次
它引用的顶会 Paper1
相关 Paper
- Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit BiasRyo Karakida, Tomoumi Takase, Tomohiro Hayase, Kazuki OsawaICML 2023 · 被引用 23 次
- iFlood: A Stable and Effective RegularizerYuexiang Xie, Zhen Wang, Yaliang Li, Ce Zhang 等ICLR 2022 · 被引用 6 次
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
- SIGUA: Forgetting May Make Learning with Noisy Labels More RobustBo Han, Gang Niu, Xingrui Yu, Quanming Yao 等ICML 2020 · 被引用 157 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
