Exploiting Invariance in Training Deep Neural Networks
Chengxi Ye, Xiong Zhou, Tristan McKinney, Yanfeng Liu, Qinggang Zhou, Fedor Zhdanov
Abstract
Inspired by two basic mechanisms in animal visual systems, we introduce a feature transform technique that imposes invariance properties in the training of deep neural networks. The resulting algorithm requires less parameter tuning, trains well with an initial learning rate 1.0, and easily generalizes to different tasks. We enforce scale invariance with local statistics in the data to align similar samples at diverse scales. To accelerate convergence, we enforce a GL(n)-invariance property with global statistics extracted from a batch such that the gradient descent solution should remain invariant under basis change. Profiling analysis shows our proposed modifications takes ∼ 5% of the computations of the underlying convolution layer. Tested on convolutional networks and transformer networks, our proposed technique requires fewer iterations to train, surpasses all baselines by a large margin, seamlessly works on both small and large batch size training, and applies to different computer vision and language tasks. In this paper, we conduct a study of a one-layer linear network to understand the origin of these limitations. Drawing inspiration from this study, the structure of animal visual systems mentioned above, and recent related work (Ye et al. 2020) , we derive two invariance properties that enhance the training of deep neural networks. We implement cross-GPU synchronization to aggregate the computation required to enforce the invariance, surpassing the widely-used synchronized batch normalization (Peng et al. 2017 ) method significantly. This implementation supports both small batch and
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20f7d703-842b-474b-baac-d730d06debb2Builds on5
- Switchable Whitening for Deep Representation LearningXingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang et al.ICCV 2019 · 204 citations
- Rethinking the Hyperparameters for Fine-tuningHao Li, Pratik Chaudhari, Hao Yang, Michael Lam et al.ICLR 2020 · 142 citations
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang et al.ICLR 2020 · 42 citations
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- Network DeconvolutionChengxi Ye, Matthew Evanusa, Hua He, Anton Mitrokhin et al.ICLR 2020
Related papers
- From Promise to Practice: Realizing High-performance Decentralized TrainingZesen Wang, Jiaojiao Zhang, Xuyang Wu, Mikael JohanssonICLR 2025
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2021 · 163 citations
- Learning in Compact Spaces with Approximately Normalized TransformerJörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev et al.NeurIPS 2025 · 5 citations
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han et al.ICLR 2021 · 165 citations
- Local Scale Equivariance with Latent Deep Equilibrium CanonicalizerMd Ashiqur Rahman, Chiao-An Yang, Michael N. Cheng, Lim Jun Hao et al.ICCV 2025
