Stochastic Normalization
Zhi Kou, Kaichao You, Mingsheng Long, Jianmin Wang
Abstract
Fine-tuning pre-trained deep networks on a small dataset is an important component in the deep learning pipeline. A critical problem in fine-tuning is how to avoid over-fitting when data are limited. Existing efforts work from two aspects: (1) impose regularization on parameters or features; (2) transfer prior knowledge to fine-tuning by reusing pre-trained parameters. In this paper, we take an alternative approach by refactoring the widely used Batch Normalization (BN) module to mitigate over-fitting. We propose a two-branch design with one branch normalized by mini-batch statistics and the other branch normalized by moving statistics. During training, two branches are stochastically selected to avoid over-depending on some sample statistics, resulting in a strong regularization effect, which we interpret as "architecture regularization." The resulting method is dubbed stochastic normalization (StochNorm). With the two-branch architecture, it naturally incorporates pre-trained moving statistics in BN layers during fine-tuning, exploiting more prior knowledge of pre-trained networks. Extensive empirical experiments show that StochNorm is a powerful tool to avoid over-fitting in fine-tuning with small datasets. Besides, StochNorm is readily pluggable in modern CNN backbones. It is complementary to other fine-tuning methods and can work together to achieve stronger regularization effect.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8372e844-481f-4e39-a9a2-1fab9e541417Cited by top-tier papers10
- Rethinking Supervised Pre-Training for Better Downstream TransferringYutong Feng, Jianwen Jiang, Mingqian Tang, Rong Jin et al.ICLR 2022 · 51 citations
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter et al.ICLR 2024 · 27 citations
- Fine-Tuning Graph Neural Networks by Preserving Graph Generative PatternsYifei Sun, Qi Zhu, Yang Yang, Chunping Wang et al.AAAI 2024 · 21 citations
- Measuring Task Similarity and Its Implication in Fine-Tuning Graph Neural NetworksRenhong Huang, Jiarong Xu, Xin Jiang, Chenglu Pan et al.AAAI 2024 · 14 citations
- Search to Fine-Tune Pre-Trained Graph Neural Networks for Graph-Level TasksZhili Wang, Shimin Di, Lei Chen, Xiaofang ZhouICDE 2024 · 6 citations
Builds on2
Related papers
- EvalNorm: Estimating Batch Normalization Statistics for EvaluationSaurabh Singh, Abhinav ShrivastavaICCV 2019 · 57 citations
- MetaNorm: Learning to Normalize Few-Shot Batches Across DomainsYing-Jun Du, Xiantong Zhen, Ling Shao, Cees G. M. SnoekICLR 2021 · 26 citations
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang et al.ICLR 2020 · 42 citations
- TWINS: A Fine-Tuning Framework for Improved Transferability of Adversarial Robustness and GeneralizationZiquan Liu, Yi Xu, Xiangyang Ji, Antoni B. ChanCVPR 2023
- Reducing Divergence in Batch Normalization for Domain AdaptationEllen Yi-Ge, Mingjing Wu, Zhenghan ChenAAAI 2025 · 3 citations
