Exploiting Invariance in Training Deep Neural Networks
Chengxi Ye, Xiong Zhou, Tristan McKinney, Yanfeng Liu, Qinggang Zhou, Fedor Zhdanov
摘要
Inspired by two basic mechanisms in animal visual systems, we introduce a feature transform technique that imposes invariance properties in the training of deep neural networks. The resulting algorithm requires less parameter tuning, trains well with an initial learning rate 1.0, and easily generalizes to different tasks. We enforce scale invariance with local statistics in the data to align similar samples at diverse scales. To accelerate convergence, we enforce a GL(n)-invariance property with global statistics extracted from a batch such that the gradient descent solution should remain invariant under basis change. Profiling analysis shows our proposed modifications takes ∼ 5% of the computations of the underlying convolution layer. Tested on convolutional networks and transformer networks, our proposed technique requires fewer iterations to train, surpasses all baselines by a large margin, seamlessly works on both small and large batch size training, and applies to different computer vision and language tasks. In this paper, we conduct a study of a one-layer linear network to understand the origin of these limitations. Drawing inspiration from this study, the structure of animal visual systems mentioned above, and recent related work (Ye et al. 2020) , we derive two invariance properties that enhance the training of deep neural networks. We implement cross-GPU synchronization to aggregate the computation required to enforce the invariance, surpassing the widely-used synchronized batch normalization (Peng et al. 2017 ) method significantly. This implementation supports both small batch and
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Switchable Whitening for Deep Representation LearningXingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang 等ICCV 2019 · 被引用 204 次
- Rethinking the Hyperparameters for Fine-tuningHao Li, Pratik Chaudhari, Hao Yang, Michael Lam 等ICLR 2020 · 被引用 142 次
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang 等ICLR 2020 · 被引用 42 次
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- Network DeconvolutionChengxi Ye, Matthew Evanusa, Hua He, Anton Mitrokhin 等ICLR 2020
相关 Paper
- From Promise to Practice: Realizing High-performance Decentralized TrainingZesen Wang, Jiaojiao Zhang, Xuyang Wu, Mikael JohanssonICLR 2025
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2021 · 被引用 163 次
- Learning in Compact Spaces with Approximately Normalized TransformerJörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev 等NeurIPS 2025 · 被引用 5 次
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han 等ICLR 2021 · 被引用 165 次
- Local Scale Equivariance with Latent Deep Equilibrium CanonicalizerMd Ashiqur Rahman, Chiao-An Yang, Michael N. Cheng, Lim Jun Hao 等ICCV 2025
