Orthogonalized SGD and Nested Architectures for Anytime Neural Networks
Chengcheng Wan, Henry Hoffmann, Shan Lu, Michael Maire
摘要
We propose a novel variant of SGD customized for training network architectures that support anytime behavior: such networks produce a series of increasingly accurate outputs over time. Efficient architectural designs for these networks focus on re-using internal state; subnetworks must produce representations relevant for both immediate prediction as well as refinement by subsequent network stages. We consider traditional branched networks as well as a new class of recursively nested networks. Our new optimizer, Orthogonalized SGD, dynamically re-balances task-specific gradients when training a multitask network. In the context of anytime architectures, this optimizer projects gradients from later outputs onto a parameter subspace that does not interfere with those from earlier outputs. Experiments demonstrate that training with Orthogonalized SGD significantly improves generalization accuracy of anytime networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Mixture of Nested Experts: Adaptive Processing of Visual TokensGagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani 等NeurIPS 2024 · 被引用 29 次
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann 等USENIX ATC 2020 · 被引用 15 次
- Accelerated Training via Incrementally Growing Neural Networks using Variance Transfer and Learning Rate AdaptationXin Yuan, Pedro Savarese, Michael MaireNeurIPS 2023 · 被引用 9 次
- Adaptive Depth Networks with Skippable Sub-PathsWoochul Kang, Hyungseop LeeNeurIPS 2024 · 被引用 5 次
相关 Paper
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 被引用 21 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu 等CVPR 2021
- Multi-Task Recurrent Modular NetworksDongkuan Xu, Wei Cheng, Xin Dong, Bo Zong 等AAAI 2021 · 被引用 2 次
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 被引用 52 次
