Toward Principled Flexible Scaling for Self-Gated Neural Activation
Sudong Cai, Shuyuan Zheng, Bingzhi Chen, Shuai Yuan, Chuan Xiao, Jianbin Qin, Bing WANG
Abstract
Neural networks necessitate nonlinearities to achieve universal approximability. Traditional activation functions introduce nonlinearities through rigid feature rectifications. Recent self-gated variants improve traditional methods in fitting flexibility by incorporating learnable content-aware factors and non-local dependencies, enabling dynamic adjustments to activation curves via adaptive translation and scaling. While SOTA approaches achieve notable gains in conventional CNN layers, they struggle to enhance Transformer layers, where fine-grained context is inherently modeled, severely reducing the effectiveness of non-local dependencies leveraged in activation processes. We refer to this critical yet unexplored challenge as the non-local tension of activation. Drawing on a decision-making perspective, we systematically analyze the origins of the non-local tension problem and explore the initial solution to foster a more discriminative and generalizable neural activation methodology. This is achieved by rethinking how non-local cues are encoded and transformed into adaptive scaling coefficients, which in turn recalibrate the contributions of features to filter updates through neural activation. Grounded in these insights, we present FleS, a novel self-gated activation model for discriminative pattern recognition. Extensive experiments on various popular benchmarks validate our interpretable methodology for improving neural activation modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ad66197-a272-43b4-857e-c6da32fecbc1Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum TechniqueKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyCVPR 2022 · 61 citations
Related papers
- AdaShift: Learning Discriminative Self-Gated Neural Feature Activation With an Adaptive Shift FactorSudong CaiCVPR 2024
- Positions, Channels, and Layers: Fully Generalized Non-Local Network for Singer IdentificationI-Yuan Kuo, Wen-Li Wei, Jen-Chun LinAAAI 2021 · 3 citations
- Unraveling Feature Extraction Mechanisms in Neural NetworksXiaobing Sun, Jiaxi Li, Wei LuEMNLP 2023
- DANet: Divergent Activation for Weakly Supervised Object LocalizationHaolan Xue, Chang Liu, Fang Wan, Jianbin Jiao et al.ICCV 2019 · 192 citations
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 10 citations
