Firefly Neural Architecture Descent: a General Approach for Growing Neural Networks
Lemeng Wu, Bo Liu, Peter Stone, Qiang Liu
Abstract
We propose firefly neural architecture descent, a general framework for progressively and dynamically growing neural networks to jointly optimize the networks' parameters and architectures. Our method works in a steepest descent fashion, which iteratively finds the best network within a functional neighborhood of the original network that includes a diverse set of candidate network structures. By using Taylor approximation, the optimal network structure in the neighborhood can be found with a greedy selection procedure. We show that firefly descent can flexibly grow networks both wider and deeper, and can be applied to learn accurate but resource-efficient neural architectures that avoid catastrophic forgetting in continual learning. Empirically, firefly descent achieves promising results on both neural architecture search and continual learning. In particular, on a challenging continual image classification task, it learns networks that are smaller in size but have higher average accuracy than those learned by the state-of-the-art methods. The code is available at https://github.com/klightz/Firefly .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 873b2ac6-e9ee-471a-beac-d89e1ead7bedCited by top-tier papers14
- GradMax: Growing Neural Networks using Gradient InformationUtku Evci, Bart van Merrienboer, Thomas Unterthiner, Fabian Pedregosa et al.ICLR 2022 · 72 citations
- Dynamically Expandable Graph Convolution for Streaming RecommendationBowei He, Xu He, Yingxue Zhang, Ruiming Tang et al.WWW 2023 · 60 citations
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingWenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang et al.NeurIPS 2024 · 52 citations
- The Elastic Lottery Ticket HypothesisXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan et al.NeurIPS 2021 · 38 citations
- Reusing Pretrained Models by Multi-linear Operators for Efficient TrainingYu Pan, Ye Yuan, Yichun Yin, Zenglin Xu et al.NeurIPS 2023 · 23 citations
Builds on1
Related papers
- CHEEM: Continual Learning by Reuse, New, Adapt and Skip - A Hierarchical Exploration-Exploitation ApproachChinmay Savadikar, Michelle Dai, Tianfu WuCVPR 2026
- Growing a Brain with Sparsity-Inducing Generation for Continual LearningHyundong Jin, Gyeong-Hyeon Kim, Chanho Ahn, Eunwoo KimICCV 2023 · 7 citations
- Meta-Continual Learning of Neural FieldsSeungyoon Woo, Junhyeog Yun, Gunhee KimICLR 2025
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 52 citations
- Growing Efficient Deep Networks by Structured Continuous SparsificationXin Yuan, Pedro Henrique Pamplona Savarese, Michael MaireICLR 2021 · 51 citations
