Stochastic Normalization
Zhi Kou, Kaichao You, Mingsheng Long, Jianmin Wang
摘要
Fine-tuning pre-trained deep networks on a small dataset is an important component in the deep learning pipeline. A critical problem in fine-tuning is how to avoid over-fitting when data are limited. Existing efforts work from two aspects: (1) impose regularization on parameters or features; (2) transfer prior knowledge to fine-tuning by reusing pre-trained parameters. In this paper, we take an alternative approach by refactoring the widely used Batch Normalization (BN) module to mitigate over-fitting. We propose a two-branch design with one branch normalized by mini-batch statistics and the other branch normalized by moving statistics. During training, two branches are stochastically selected to avoid over-depending on some sample statistics, resulting in a strong regularization effect, which we interpret as "architecture regularization." The resulting method is dubbed stochastic normalization (StochNorm). With the two-branch architecture, it naturally incorporates pre-trained moving statistics in BN layers during fine-tuning, exploiting more prior knowledge of pre-trained networks. Extensive empirical experiments show that StochNorm is a powerful tool to avoid over-fitting in fine-tuning with small datasets. Besides, StochNorm is readily pluggable in modern CNN backbones. It is complementary to other fine-tuning methods and can work together to achieve stronger regularization effect.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Rethinking Supervised Pre-Training for Better Downstream TransferringYutong Feng, Jianwen Jiang, Mingqian Tang, Rong Jin 等ICLR 2022 · 被引用 51 次
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter 等ICLR 2024 · 被引用 27 次
- Fine-Tuning Graph Neural Networks by Preserving Graph Generative PatternsYifei Sun, Qi Zhu, Yang Yang, Chunping Wang 等AAAI 2024 · 被引用 21 次
- Measuring Task Similarity and Its Implication in Fine-Tuning Graph Neural NetworksRenhong Huang, Jiarong Xu, Xin Jiang, Chenglu Pan 等AAAI 2024 · 被引用 14 次
- Search to Fine-Tune Pre-Trained Graph Neural Networks for Graph-Level TasksZhili Wang, Shimin Di, Lei Chen, Xiaofang ZhouICDE 2024 · 被引用 6 次
它引用的顶会 Paper2
相关 Paper
- EvalNorm: Estimating Batch Normalization Statistics for EvaluationSaurabh Singh, Abhinav ShrivastavaICCV 2019 · 被引用 57 次
- MetaNorm: Learning to Normalize Few-Shot Batches Across DomainsYing-Jun Du, Xiantong Zhen, Ling Shao, Cees G. M. SnoekICLR 2021 · 被引用 26 次
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang 等ICLR 2020 · 被引用 42 次
- TWINS: A Fine-Tuning Framework for Improved Transferability of Adversarial Robustness and GeneralizationZiquan Liu, Yi Xu, Xiangyang Ji, Antoni B. ChanCVPR 2023
- Reducing Divergence in Batch Normalization for Domain AdaptationEllen Yi-Ge, Mingjing Wu, Zhenghan ChenAAAI 2025 · 被引用 3 次
