AdaShift: Learning Discriminative Self-Gated Neural Feature Activation With an Adaptive Shift Factor
Sudong Cai
Abstract
Nonlinearities are decisive in neural representation learning. Traditional Activation (Act) functions impose fixed inductive biases on neural networks with oriented biological intuitions. Recent methods leverage selfgated curves to compensate for the rigid traditional Act paradigms in fitting flexibility. However, substantial improvements are still impeded by the norm-induced mismatched feature re-calibrations (see Section 1), i.e., the actual importance of a feature can be inconsistent with its explicit intensity such that violates the basic intention of a direct self-gated feature re-weighting. To address this problem, we propose to learn discriminative neural feature Act with a novel prototype, namely, AdaShift, which enhances typical self-gated Act by incorporating an adaptive shift factor into the re-weighting function of Act. AdaShift casts dynamic translations on the inputs of a re-weighting function by exploiting comprehensive feature-filter context cues of different ranges in a simple yet effective manner. We obtain the new intuitions of AdaShift by rethinking the feature-filter relationships from a common Softmax-based classification and by generalizing the new observations to a common learning layer that encodes features with updatable filters. Our practical AdaShifts, built upon the new Act prototype, demonstrate significant improvements to the popular/SOTA Act functions on different vision benchmarks. By simply replacing ReLU with AdaShifts, ResNets can match advanced Transformer counterparts (e.g., ResNet-50 vs. Swin-T) with lower cost and fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99c9dcb1-4a2d-43f9-babf-21258f587f28Builds on11
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Ground-to-Aerial Image Geo-Localization With a Hard Exemplar Reweighting Triplet LossSudong Cai, Yulan Guo, Salman H. Khan, Jiwei Hu et al.ICCV 2019 · 140 citations
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep NetworksAlejandro Molina, Patrick Schramowski, Kristian KerstingICLR 2020 · 116 citations
- Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum TechniqueKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyCVPR 2022 · 61 citations
Related papers
- IIEU: Rethinking Neural Feature Activation from Decision-MakingSudong CaiICCV 2023 · 1 citation
- Stochastic Adaptive Activation FunctionKyungsu Lee, Jaeseung Yang, Haeyun Lee, Jae Youn HwangNeurIPS 2022 · 6 citations
- Toward Principled Flexible Scaling for Self-Gated Neural ActivationSudong Cai, Shuyuan Zheng, Bingzhi Chen, Shuai Yuan et al.ICLR 2026
- AcTTA: Rethinking Test-Time Adaptation via Dynamic ActivationHyeongyu Kim, Geonhui Han, Dosik HwangCVPR 2026
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 10 citations
