TaskNorm: Rethinking Batch Normalization for Meta-Learning
John Bronskill, Jonathan Gordon, James Requeima, Sebastian Nowozin, Richard E. Turner
Abstract
Modern meta-learning approaches for image classification rely on increasingly deep networks to achieve state-of-the-art performance, making batch normalization an essential component of meta-learning pipelines. However, the hierarchical nature of the meta-learning setting presents several challenges that can render conventional batch normalization ineffective, giving rise to the need to rethink normalization in this setting. We evaluate a range of approaches to batch normalization for meta-learning scenarios, and develop a novel approach that we call TaskNorm. Experiments on fourteen datasets demonstrate that the choice of batch normalization has a dramatic effect on both classification accuracy and training time for both gradient based and gradient-free meta-learning approaches. Importantly, TaskNorm is found to consistently improve performance. Finally, we provide a set of best practices for normalization that will allow fair comparison of meta-learning algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7bc425f5-c442-4257-a6bb-dbf7191b1e2eCited by top-tier papers22
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang et al.ICLR 2021 · 143 citations
- Universal Representation Learning from Multiple Domains for Few-shot ClassificationWei-Hong Li, Xialei Liu, Hakan BilenICCV 2021 · 114 citations
- Continual Normalization: Rethinking Batch Normalization for Online Continual LearningQuang Pham, Chenghao Liu, Steven C. H. HoiICLR 2022 · 72 citations
- Few-Shot One-Class Classification via Meta-LearningAhmed Frikha, Denis Krompaß, Hans-Georg Köpken, Volker TrespAAAI 2021 · 68 citations
- Learning where to learn: Gradient sparsity in meta and continual learningJohannes von Oswald, Dominic Zhao, Seijin Kobayashi, Simon Schug et al.NeurIPS 2021 · 61 citations
Builds on2
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
Related papers
- MetaNorm: Learning to Normalize Few-Shot Batches Across DomainsYing-Jun Du, Xiantong Zhen, Ling Shao, Cees G. M. SnoekICLR 2021 · 26 citations
- MetaModulation: Learning Variational Feature Hierarchies for Few-Shot Learning with Fewer TasksWenfang Sun, Yingjun Du, Xiantong Zhen, Fan Wang et al.ICML 2023 · 10 citations
- Meta Batch-Instance Normalization for Generalizable Person Re-IdentificationSeokeon Choi, Taekyung Kim, Minki Jeong, Hyoungseob Park et al.CVPR 2021
- Meta-Learning without MemorizationMingzhang Yin, George Tucker, Mingyuan Zhou, Sergey Levine et al.ICLR 2020 · 201 citations
- Is normalization indispensable for training deep neural network?Jie Shao, Kai Hu, Changhu Wang, Xiangyang Xue et al.NeurIPS 2020 · 70 citations
