Evolving Normalization-Activation Layers
Hanxiao Liu, Andy Brock, Karen Simonyan, Quoc Le
摘要
Normalization layers and activation functions are fundamental components in deep networks and typically co-locate with each other. Here we propose to design them using an automated approach. Instead of designing them separately, we unify them into a single tensor-to-tensor computation graph, and evolve its structure starting from basic mathematical functions. Examples of such mathematical functions are addition, multiplication and statistical moments. The use of low-level mathematical functions, in contrast to the use of high-level modules in mainstream NAS, leads to a highly sparse and large search space which can be challenging for search methods. To address the challenge, we develop efficient rejection protocols to quickly filter out candidate layers that do not work well. We also use multiobjective evolution to optimize each layer's performance across many architectures to prevent overfitting. Our method leads to the discovery of EvoNorms, a set of new normalization-activation layers with novel, and sometimes surprising structures that go beyond existing design patterns. For example, some EvoNorms do not assume that normalization and activation functions must be applied sequentially, nor need to center the feature maps, nor require explicit activation functions. Our experiments show that EvoNorms work well on image classification models including ResNets, MobileNets and EfficientNets but also transfer well to Mask R-CNN with FPN/SpineNet for instance segmentation and to BigGAN for image synthesis, outperforming BatchNorm and GroupNorm based layers in many cases. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Searching for Efficient Transformers for Language ModelingDavid R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai 等NeurIPS 2021 · 被引用 205 次
- EvoPrompting: Language Models for Code-Level Neural Architecture SearchAngelica Chen, David Dohan, David R. SoNeurIPS 2023 · 被引用 184 次
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
- RelaySum for Decentralized Deep Learning on Heterogeneous DataThijs Vogels, Lie He, Anastasia Koloskova, Sai Praneeth Karimireddy 等NeurIPS 2021 · 被引用 78 次
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 被引用 50 次
它引用的顶会 Paper7
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 被引用 725 次
- Exploring Randomly Wired Neural Networks for Image RecognitionSaining Xie, Alexander Kirillov, Ross B. Girshick, Kaiming HeICCV 2019 · 被引用 384 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
相关 Paper
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- AutoML-Zero: Evolving Machine Learning Algorithms From ScratchEsteban Real, Chen Liang, David R. So, Quoc V. LeICML 2020 · 被引用 265 次
- Divergence-Free Neural Networks with Application to Image DenoisingSébastien Herbreteau, Etienne MeunierICLR 2026
- Equivariance-aware Architectural Optimization of Neural NetworksKaitlin Maile, Dennis George Wilson, Patrick ForréICLR 2023
- MUXConv: Information Multiplexing in Convolutional Neural NetworksZhichao Lu, Kalyanmoy Deb, Vishnu Naresh BoddetiCVPR 2020
