Lune

AAAI2022顶会

Exploiting Invariance in Training Deep Neural Networks

Chengxi Ye, Xiong Zhou, Tristan McKinney, Yanfeng Liu, Qinggang Zhou, Fedor Zhdanov

2022年份
4被引次数

摘要

Inspired by two basic mechanisms in animal visual systems, we introduce a feature transform technique that imposes invariance properties in the training of deep neural networks. The resulting algorithm requires less parameter tuning, trains well with an initial learning rate 1.0, and easily generalizes to different tasks. We enforce scale invariance with local statistics in the data to align similar samples at diverse scales. To accelerate convergence, we enforce a GL(n)-invariance property with global statistics extracted from a batch such that the gradient descent solution should remain invariant under basis change. Profiling analysis shows our proposed modifications takes ∼ 5% of the computations of the underlying convolution layer. Tested on convolutional networks and transformer networks, our proposed technique requires fewer iterations to train, surpasses all baselines by a large margin, seamlessly works on both small and large batch size training, and applies to different computer vision and language tasks. In this paper, we conduct a study of a one-layer linear network to understand the origin of these limitations. Drawing inspiration from this study, the structure of animal visual systems mentioned above, and recent related work (Ye et al. 2020) , we derive two invariance properties that enhance the training of deep neural networks. We implement cross-GPU synchronization to aggregate the computation required to enforce the invariance, surpassing the widely-used synchronized batch normalization (Peng et al. 2017 ) method significantly. This implementation supports both small batch and

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖