WLD-Reg: A Data-Dependent Within-Layer Diversity Regularizer
Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj
Abstract
Neural networks are composed of multiple layers arranged in a hierarchical structure jointly trained with a gradient-based optimization, where the errors are back-propagated from the last layer back to the first one. At each optimization step, neurons at a given layer receive feedback from neurons belonging to higher layers of the hierarchy. In this paper, we propose to complement this traditional 'between-layer' feedback with additional 'within-layer' feedback to encourage the diversity of the activations within the same layer. To this end, we measure the pairwise similarity between the outputs of the neurons and use it to model the layer's overall diversity. We present an extensive empirical study confirming that the proposed approach enhances the performance of several state-of-the-art neural network models in multiple tasks. The code is publically available at https://github.com/firasl/AAAI-23-WLD-Reg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74ba3470-cb8a-4879-8146-9d05962fdf82Cited by top-tier papers2
- Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal FeaturesAnnie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine et al.ICLR 2024 · 5 citations
- Test-time Diverse Reasoning by Riemannian Activation SteeringLy Tran Ho Khanh, Dongxuan Zhu, Man-Chung Yue, Viet Anh NguyenAAAI 2026 · 1 citation
Builds on6
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 912 citations
Related papers
- Lamina-specific neuronal properties promote robust, stable signal propagation in feedforward networksDongqi Han, Erik De Schutter, Sungho HongNeurIPS 2020 · 3 citations
- MMA Regularization: Decorrelating Weights of Neural Networks by Maximizing the Minimal AnglesZhennan Wang, Canqun Xiang, Wenbin Zou, Chen XuNeurIPS 2020 · 25 citations
- Neuron with Steady Response Leads to Better GeneralizationQiang Fu, Lun Du, Haitao Mao, Xu Chen et al.NeurIPS 2022 · 5 citations
- Fixed-Weight Difference Target PropagationTatsukichi Shibuya, Nakamasa Inoue, Rei Kawakami, Ikuro SatoAAAI 2023 · 6 citations
- Intraclass clustering: an implicit learning ability that regularizes DNNsSimon Carbonnelle, Christophe De VleeschouwerICLR 2021 · 2 citations
