Dynamic Hierarchical Mimicking Towards Consistent Optimization Objectives
Duo Li, Qifeng Chen
摘要
While the depth of modern Convolutional Neural Networks (CNNs) surpasses that of the pioneering networks with a significant margin, the traditional way of appending supervision only over the final classifier and progressively propagating gradient flow upstream remains the training mainstay. Seminal Deeply-Supervised Networks (DSN) were proposed to alleviate the difficulty of optimization arising from gradient flow through a long chain. However, it is still vulnerable to issues including interference to the hierarchical representation generation process and inconsistent optimization objectives, as illustrated theoretically and empirically in this paper. Complementary to previous training strategies, we propose Dynamic Hierarchical Mimicking, a generic feature learning mechanism, to advance CNN training with enhanced generalization ability. Partially inspired by DSN, we fork delicately designed side branches from the intermediate layers of a given neural network. Each branch can emerge from certain locations of the main branch dynamically, which not only retains representation rooted in the backbone network but also generates more diverse representations along its own pathway. We go one step further to promote multi-level interactions among different branches through an optimization formula with probabilistic prediction matching losses, thus guaranteeing a more robust optimization process and better representation ability. Experiments on both category and instance recognition tasks demonstrate the substantial improvements of our proposed method over its corresponding counterparts using diverse state-of-the-art CNN architectures. Code and models are publicly available at https://github.com/d-li14/DHM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Differentiable Dynamic Wirings for Neural NetworksKun Yuan, Quanquan Li, Shaopeng Guo, Dapeng Chen 等ICCV 2021 · 被引用 8 次
- Group-wise Inhibition based Feature Regularization for Robust ClassificationHaozhe Liu, Haoqian Wu, Weicheng Xie, Feng Liu 等ICCV 2021 · 被引用 17 次
- Multi-Proxy Wasserstein Classifier for Image ClassificationBenlin Liu, Yongming Rao, Jiwen Lu, Jie Zhou 等AAAI 2021 · 被引用 10 次
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang 等ICML 2023 · 被引用 60 次
- IA-GM: A Deep Bidirectional Learning Method for Graph MatchingKaixuan Zhao, Shikui Tu, Lei XuAAAI 2021 · 被引用 13 次
