Learning Compact Features via In-Training Representation Alignment
Xin Li, Xiangrui Li, Deng Pan, Yao Qiang, Dongxiao Zhu
Abstract
Deep neural networks (DNNs) for supervised learning can be viewed as a pipeline of the feature extractor (i.e., last hidden layer) and a linear classifier (i.e., output layer) that are trained jointly with stochastic gradient descent (SGD) on the loss function (e.g., cross-entropy). In each epoch, the true gradient of the loss function is estimated using a mini-batch sampled from the training set and model parameters are then updated with the mini-batch gradients. Although the latter provides an unbiased estimation of the former, they are subject to substantial variances derived from the size and number of sampled mini-batches, leading to noisy and jumpy updates. To stabilize such undesirable variance in estimating the true gradients, we propose In-Training Representation Alignment (ITRA) that explicitly aligns feature distributions of two different mini-batches with a matching loss in the SGD training process. We also provide a rigorous analysis of the desirable effects of the matching loss on feature representation learning: (1) extracting compact feature representation; (2) reducing over-adaption on mini-batches via an adaptively weighting mechanism; and (3) accommodating to multi-modalities. Finally, we conduct large-scale experiments on both image and text classifications to demonstrate its superior performance to the strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfb550f1-37b3-4fbc-9797-9cf7e227e02fCited by top-tier papers1
Ask how each one uses itBuilds on3
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- On the Learning Property of Logistic and Softmax Losses for Deep Neural NetworksXiangrui Li, Xin Li, Deng Pan, Dongxiao ZhuAAAI 2020 · 26 citations
- Improving Adversarial Robustness via Probabilistically Compact Loss with Logit ConstraintsXin Li, Xiangrui Li, Deng Pan, Dongxiao ZhuAAAI 2021 · 17 citations
Related papers
- Adaptive Random Feature Regularization on Fine-tuning Deep Neural NetworksShin'ya Yamaguchi, Sekitoshi Kanai, Kazuki Adachi, Daiki ChijiwaCVPR 2024 · 3 citations
- Procrustean Training for Imbalanced Deep LearningHan-Jia Ye, De-Chuan Zhan, Wei-Lun ChaoICCV 2021 · 36 citations
- Maintaining Discrimination and Fairness in Class Incremental LearningBowen Zhao, Xi Xiao, Guojun Gan, Bin Zhang et al.CVPR 2020
- Why Do Better Loss Functions Lead to Less Transferable Features?Simon Kornblith, Ting Chen, Honglak Lee, Mohammad NorouziNeurIPS 2021 · 113 citations
- Align, then memorise: the dynamics of learning with feedback alignmentMaria Refinetti, Stéphane d'Ascoli, Ruben Ohana, Sebastian GoldtICML 2021 · 47 citations
