Stimulative Training of Residual Networks: A Social Psychology Perspective of Loafing
Peng Ye, Shengji Tang, Baopu Li, Tao Chen, Wanli Ouyang
Abstract
Residual networks have shown great success and become indispensable in today's deep models. In this work, we aim to re-investigate the training process of residual networks from a novel social psychology perspective of loafing, and further propose a new training strategy to strengthen the performance of residual networks. As residual networks can be viewed as ensembles of relatively shallow networks (i.e., unraveled view) in prior works, we also start from such view and consider that the final performance of a residual network is co-determined by a group of sub-networks. Inspired by the social loafing problem of social psychology, we find that residual networks invariably suffer from similar problem, where sub-networks in a residual network are prone to exert less effort when working as part of the group compared to working alone. We define this previously overlooked problem as network loafing. As social loafing will ultimately cause the low individual productivity and the reduced overall performance, network loafing will also hinder the performance of a given residual network and its sub-networks. Referring to the solutions of social psychology, we propose stimulative training, which randomly samples a residual sub-network and calculates the KL-divergence loss between the sampled sub-network and the given residual network, to act as extra supervision for sub-networks and make the overall goal consistent. Comprehensive empirical results and theoretical analyses verify that stimulative training can well handle the loafing problem, and improve the performance of a residual network by improving the performance of its sub-networks. The code is available at https://github.com/Sunshine-Ye/NIPS22-ST .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- EMR-Merging: Tuning-Free High-Performance Model MergingChenyu Huang, Peng Ye, Tao Chen, Tong He et al.NeurIPS 2024 · 134 citations
- Boosting Residual Networks with Group KnowledgeShengji Tang, Peng Ye, Baopu Li, Weihao Lin et al.AAAI 2024 · 7 citations
- S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in PruningWeihao Lin, Shengji Tang, Chong Yu, Peng Ye et al.NeurIPS 2024 · 2 citations
- PaceLLM: Brain-Inspired Large Language Models for Long-Context UnderstandingKangcong Li, Peng Ye, Chongjun Tu, Lin Zhang et al.NeurIPS 2025
Builds on8
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- CompOFA - Compound Once-For-All Networks for Faster Multi-Platform DeploymentManas Sahni, Shreya Varshini, Alind Khare, Alexey TumanovICLR 2021 · 37 citations
Related papers
- Towards Adaptive Residual Network Training: A Neural-ODE PerspectiveChengyu Dong, Liyuan Liu, Zichao Li, Jingbo ShangICML 2020 · 35 citations
- Co-training 2L Submodels for Visual RecognitionHugo Touvron, Matthieu Cord, Maxime Oquab, Piotr Bojanowski et al.CVPR 2023
- Efficient stagewise pretraining via progressive subnetworksAbhishek Panigrahi, Nikunj Saunshi, Kaifeng Lyu, Sobhan Miryoosefi et al.ICLR 2025
- Surrogate Module Learning: Reduce the Gradient Error Accumulation in Training Spiking Neural NetworksShikuang Deng, Hao Lin, Yuhang Li, Shi GuICML 2023 · 36 citations
- Regularization in ResNet with Stochastic DepthSoufiane Hayou, Fadhel AyedNeurIPS 2021 · 17 citations
