Progressive Neural Architecture Generation
Caiyang Yu, Chen Huang, Yun Liu, Chenwei Tang, Wei Ju, Jiancheng Lv
Abstract
Specifically, MSQ constructs sub-architectures using quantization decoding and progressively expands them, transitioning from simple to complex forms. This operation bypasses network inference to enhance efficiency. Complementing MSQ, SCC, implemented through a tailored regularization mechanism, introduces penalties for deviations during subarchitecture generation, guiding the process towards valid target architectures. As such, PNAG establishes a clear generation path, laying the groundwork for generating suitable architectures in downstream tasks. Extensive experiments demonstrate that PNAG not only generates superior architectures for various downstream tasks (+8.43%/+5.07%, on average) but also significantly improves generation efficiency, reducing the architecture generation time by 1300×. Furthermore, PNAG demonstrates strong extensibility by successfully generating Transformer-based architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 352c6db9-0190-47e9-b9ef-4fee6901fab5Builds on21
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen et al.ICLR 2020 · 691 citations
Related papers
- DiffusionNAG: Predictor-guided Neural Architecture Generation with Diffusion ModelsSohyun An, Hayeon Lee, Jaehyeong Jo, Seanie Lee et al.ICLR 2024 · 21 citations
- Towards Next-Level Post-Training Quantization of Hyper-Scale TransformersJunhan Kim, Chungman Lee, Eulrang Cho, Kyungphil Park et al.NeurIPS 2024 · 10 citations
- CSTrans-OPU: An FPGA-based Overlay Processor with Full Compilation for Transformer Networks via Sparsity ExplorationYueyin Bai, Keqing Zhao, Yang Liu, Hongji Wang et al.DAC 2024 · 4 citations
- Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive NetworksEeshaan Jain, Tushar Nandy, Gaurav Aggarwal, Ashish Tendulkar et al.NeurIPS 2023 · 30 citations
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference AccelerationJintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu et al.ICLR 2025
