Orthogonalized SGD and Nested Architectures for Anytime Neural Networks
Chengcheng Wan, Henry Hoffmann, Shan Lu, Michael Maire
Abstract
We propose a novel variant of SGD customized for training network architectures that support anytime behavior: such networks produce a series of increasingly accurate outputs over time. Efficient architectural designs for these networks focus on re-using internal state; subnetworks must produce representations relevant for both immediate prediction as well as refinement by subsequent network stages. We consider traditional branched networks as well as a new class of recursively nested networks. Our new optimizer, Orthogonalized SGD, dynamically re-balances task-specific gradients when training a multitask network. In the context of anytime architectures, this optimizer projects gradients from later outputs onto a parameter subspace that does not interfere with those from earlier outputs. Experiments demonstrate that training with Orthogonalized SGD significantly improves generalization accuracy of anytime networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88f2cf06-fcce-45b2-9d72-807c7c44740eCited by top-tier papers4
- Mixture of Nested Experts: Adaptive Processing of Visual TokensGagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani et al.NeurIPS 2024 · 29 citations
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann et al.USENIX ATC 2020 · 15 citations
- Accelerated Training via Incrementally Growing Neural Networks using Variance Transfer and Learning Rate AdaptationXin Yuan, Pedro Savarese, Michael MaireNeurIPS 2023 · 9 citations
- Adaptive Depth Networks with Skippable Sub-PathsWoochul Kang, Hyungseop LeeNeurIPS 2024 · 5 citations
Related papers
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 21 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu et al.CVPR 2021
- Multi-Task Recurrent Modular NetworksDongkuan Xu, Wei Cheng, Xin Dong, Bo Zong et al.AAAI 2021 · 2 citations
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 52 citations
