Usable Information and Evolution of Optimal Representations During Training
Michael Kleinman, Alessandro Achille, Daksh Idnani, Jonathan C. Kao
Abstract
We introduce a notion of usable information contained in the representation learned by a deep network, and use it to study how optimal representations for the task emerge during training. We show that the implicit regularization coming from training with Stochastic Gradient Descent with a high learning-rate and small batch size plays an important role in learning minimal sufficient representations for the task. In the process of arriving at a minimal sufficient representation, we find that the content of the representation changes dynamically during training. In particular, we find that semantically meaningful but ultimately irrelevant information is encoded in the early transient dynamics of training, before being later discarded. In addition, we evaluate how perturbing the initial part of training impacts the learning dynamics and the resulting representations. We show these effects on both perceptual decision-making tasks inspired by neuroscience literature, as well as on standard image classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 453a01fe-f58b-4ace-aa7d-d13f1898e70aCited by top-tier papers5
- Gacs-Korner Common Information Variational AutoencoderMichael Kleinman, Alessandro Achille, Stefano Soatto, Jonathan C. KaoNeurIPS 2023 · 23 citations
- A mechanistic multi-area recurrent network model of decision-makingMichael Kleinman, Chandramouli Chandrasekaran, Jonathan C. KaoNeurIPS 2021 · 19 citations
- Does YOLO Really Need to See Every Training Image in Every Epoch?Xingxing Xie, Jiahua Dong, Junwei Han, Gong ChengCVPR 2026 · 1 citation
- Understanding the Learning Phases in Self-Supervised Learning via Critical PeriodsJanghyeon Lee, Philipe A. Dias, Yao-Yi Chiang, Dalton D. LungaICLR 2026
- Critical Learning Periods for Multisensory Integration in Deep NetworksMichael Kleinman, Alessandro Achille, Stefano SoattoCVPR 2023
Builds on1
Related papers
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 77 citations
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
- To Grok or not to Grok: Disentangling Generalization and Memorization on Corrupted Algorithmic DatasetsDarshil Doshi, Aritra Das, Tianyu He, Andrey GromovICLR 2024 · 23 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg et al.ICML 2021 · 78 citations
