Co-training 2L Submodels for Visual Recognition
Hugo Touvron, Matthieu Cord, Maxime Oquab, Piotr Bojanowski, Jakob Verbeek, Hervé Jégou
摘要
We introduce submodel co-training, a regularization method related to co-training, self-distillation and stochastic depth. Given a neural network to be trained, for each sample we implicitly instantiate two altered networks, "submodels", with stochastic depth: we activate only a subset of the layers. Each network serves as a soft teacher to the other, by providing a loss that complements the regular loss provided by the one-hot label. Our approach, dubbed "cosub", uses a single set of weights, and does not involve a pre-trained external model or temporal averaging. Experimentally, we show that submodel co-training is effective to train backbones for recognition tasks such as image classification and semantic segmentation. Our approach is compatible with multiple architectures, including
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Switching Temporary Teachers for Semi-Supervised Semantic SegmentationJaemin Na, Jung-Woo Ha, Hyung Jin Chang, Dongyoon Han 等NeurIPS 2023 · 被引用 72 次
- Vision-LSTM: xLSTM as Generic Vision BackboneBenedikt Alkin, Maximilian Beck, Korbinian Pöppel, Sepp Hochreiter 等ICLR 2025 · 被引用 20 次
- Adaptive Depth Networks with Skippable Sub-PathsWoochul Kang, Hyungseop LeeNeurIPS 2024 · 被引用 5 次
- Sparse Cocktail: Every Sparse Pattern Every Sparse Ratio All At OnceZhangheng Li, Shiwei Liu, Tianlong Chen, Ajay Kumar Jaiswal 等ICML 2024 · 被引用 2 次
- CR-CTC: Consistency regularization on CTC for improved speech recognitionZengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang 等ICLR 2025
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff PerspectiveHelong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou 等ICLR 2021 · 被引用 209 次
- Boosting Residual Networks with Group KnowledgeShengji Tang, Peng Ye, Baopu Li, Weihao Lin 等AAAI 2024 · 被引用 7 次
- Network as Regularization for Training Deep Neural Networks: Framework, Model and PerformanceKai Tian, Yi Xu, Jihong Guan, Shuigeng ZhouAAAI 2020 · 被引用 6 次
- Regularization in ResNet with Stochastic DepthSoufiane Hayou, Fadhel AyedNeurIPS 2021 · 被引用 17 次
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
