Progressive Neural Architecture Generation
Caiyang Yu, Chen Huang, Yun Liu, Chenwei Tang, Wei Ju, Jiancheng Lv
摘要
Specifically, MSQ constructs sub-architectures using quantization decoding and progressively expands them, transitioning from simple to complex forms. This operation bypasses network inference to enhance efficiency. Complementing MSQ, SCC, implemented through a tailored regularization mechanism, introduces penalties for deviations during subarchitecture generation, guiding the process towards valid target architectures. As such, PNAG establishes a clear generation path, laying the groundwork for generating suitable architectures in downstream tasks. Extensive experiments demonstrate that PNAG not only generates superior architectures for various downstream tasks (+8.43%/+5.07%, on average) but also significantly improves generation efficiency, reducing the architecture generation time by 1300×. Furthermore, PNAG demonstrates strong extensibility by successfully generating Transformer-based architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen 等ICLR 2020 · 被引用 691 次
相关 Paper
- DiffusionNAG: Predictor-guided Neural Architecture Generation with Diffusion ModelsSohyun An, Hayeon Lee, Jaehyeong Jo, Seanie Lee 等ICLR 2024 · 被引用 21 次
- Towards Next-Level Post-Training Quantization of Hyper-Scale TransformersJunhan Kim, Chungman Lee, Eulrang Cho, Kyungphil Park 等NeurIPS 2024 · 被引用 10 次
- CSTrans-OPU: An FPGA-based Overlay Processor with Full Compilation for Transformer Networks via Sparsity ExplorationYueyin Bai, Keqing Zhao, Yang Liu, Hongji Wang 等DAC 2024 · 被引用 4 次
- Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive NetworksEeshaan Jain, Tushar Nandy, Gaurav Aggarwal, Ashish Tendulkar 等NeurIPS 2023 · 被引用 30 次
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference AccelerationJintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu 等ICLR 2025
