Improving Gradient Flow with Unrolled Highway Expectation Maximization
Chonghyuk Song, Eunseok Kim, Inwook Shim
摘要
Integrating model-based machine learning methods into deep neural architectures allows one to leverage both the expressive power of deep neural nets and the ability of model-based methods to incorporate domain-specific knowledge. In particular, many works have employed the expectation maximization (EM) algorithm in the form of an unrolled layer-wise structure that is jointly trained with a backbone neural network. However, it is difficult to discriminatively train the backbone network by backpropagating through the EM iterations as they are prone to the vanishing gradient problem. To address this issue, we propose Highway Expectation Maximization Networks (HEMNet), which is comprised of unrolled iterations of the generalized EM (GEM) algorithm based on the Newton-Rahpson method. HEMNet features scaled skip connections, or highways, along the depths of the unrolled architecture, resulting in improved gradient flow during backpropagation while incurring negligible additional computation and memory costs compared to standard unrolled EM. Furthermore, HEMNet preserves the underlying EM procedure, thereby fully retaining the convergence properties of the original EM algorithm. We achieve significant improvement in performance on several semantic segmentation benchmarks and empirically show that HEMNet effectively alleviates gradient decay.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang 等ICCV 2019 · 被引用 639 次
- Weakly Supervised Fine-Grained Image Classification via Guassian Mixture Model Oriented Discriminative LearningZhihui Wang, Shijie Wang, Shuhui Yang, Haojie Li 等CVPR 2020
相关 Paper
- Highway Value Iteration NetworksYuhui Wang, Weida Li, Francesco Faccio, Qingyuan Wu 等ICML 2024 · 被引用 3 次
- Accelerated training through iterative gradient propagation along the residual pathErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICLR 2025
- An EM Framework for Online Incremental Learning of Semantic SegmentationShipeng Yan, Jiale Zhou, Jiangwei Xie, Songyang Zhang 等ACM MM 2021 · 被引用 32 次
- CNM-UNet: Continuous Ordinary Differential Equations for Medical Image SegmentationTianqi Xu, Yashi Zhu, Quansong He, Yue Cao 等AAAI 2026
- Beyond Skip Connection: Pooling and Unpooling Design for Elimination SingularitiesChengkun Sun, Jinqian Pan, Zhuoli Jin, Russell Stevens Terry 等AAAI 2025 · 被引用 1 次
