Firefly Neural Architecture Descent: a General Approach for Growing Neural Networks
Lemeng Wu, Bo Liu, Peter Stone, Qiang Liu
摘要
We propose firefly neural architecture descent, a general framework for progressively and dynamically growing neural networks to jointly optimize the networks' parameters and architectures. Our method works in a steepest descent fashion, which iteratively finds the best network within a functional neighborhood of the original network that includes a diverse set of candidate network structures. By using Taylor approximation, the optimal network structure in the neighborhood can be found with a greedy selection procedure. We show that firefly descent can flexibly grow networks both wider and deeper, and can be applied to learn accurate but resource-efficient neural architectures that avoid catastrophic forgetting in continual learning. Empirically, firefly descent achieves promising results on both neural architecture search and continual learning. In particular, on a challenging continual image classification task, it learns networks that are smaller in size but have higher average accuracy than those learned by the state-of-the-art methods. The code is available at https://github.com/klightz/Firefly .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- GradMax: Growing Neural Networks using Gradient InformationUtku Evci, Bart van Merrienboer, Thomas Unterthiner, Fabian Pedregosa 等ICLR 2022 · 被引用 72 次
- Dynamically Expandable Graph Convolution for Streaming RecommendationBowei He, Xu He, Yingxue Zhang, Ruiming Tang 等WWW 2023 · 被引用 60 次
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingWenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang 等NeurIPS 2024 · 被引用 52 次
- The Elastic Lottery Ticket HypothesisXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan 等NeurIPS 2021 · 被引用 38 次
- Reusing Pretrained Models by Multi-linear Operators for Efficient TrainingYu Pan, Ye Yuan, Yichun Yin, Zenglin Xu 等NeurIPS 2023 · 被引用 23 次
它引用的顶会 Paper1
相关 Paper
- CHEEM: Continual Learning by Reuse, New, Adapt and Skip - A Hierarchical Exploration-Exploitation ApproachChinmay Savadikar, Michelle Dai, Tianfu WuCVPR 2026
- Growing a Brain with Sparsity-Inducing Generation for Continual LearningHyundong Jin, Gyeong-Hyeon Kim, Chanho Ahn, Eunwoo KimICCV 2023 · 被引用 7 次
- Meta-Continual Learning of Neural FieldsSeungyoon Woo, Junhyeog Yun, Gunhee KimICLR 2025
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 被引用 52 次
- Growing Efficient Deep Networks by Structured Continuous SparsificationXin Yuan, Pedro Henrique Pamplona Savarese, Michael MaireICLR 2021 · 被引用 51 次
