Improving Gradient Flow with Unrolled Highway Expectation Maximization
Chonghyuk Song, Eunseok Kim, Inwook Shim
Abstract
Integrating model-based machine learning methods into deep neural architectures allows one to leverage both the expressive power of deep neural nets and the ability of model-based methods to incorporate domain-specific knowledge. In particular, many works have employed the expectation maximization (EM) algorithm in the form of an unrolled layer-wise structure that is jointly trained with a backbone neural network. However, it is difficult to discriminatively train the backbone network by backpropagating through the EM iterations as they are prone to the vanishing gradient problem. To address this issue, we propose Highway Expectation Maximization Networks (HEMNet), which is comprised of unrolled iterations of the generalized EM (GEM) algorithm based on the Newton-Rahpson method. HEMNet features scaled skip connections, or highways, along the depths of the unrolled architecture, resulting in improved gradient flow during backpropagation while incurring negligible additional computation and memory costs compared to standard unrolled EM. Furthermore, HEMNet preserves the underlying EM procedure, thereby fully retaining the convergence properties of the original EM algorithm. We achieve significant improvement in performance on several semantic segmentation benchmarks and empirically show that HEMNet effectively alleviates gradient decay.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa119516-1a8e-441a-aca4-4d758133bd0bCited by top-tier papers1
Ask how each one uses itBuilds on2
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang et al.ICCV 2019 · 639 citations
- Weakly Supervised Fine-Grained Image Classification via Guassian Mixture Model Oriented Discriminative LearningZhihui Wang, Shijie Wang, Shuhui Yang, Haojie Li et al.CVPR 2020
Related papers
- Highway Value Iteration NetworksYuhui Wang, Weida Li, Francesco Faccio, Qingyuan Wu et al.ICML 2024 · 3 citations
- Accelerated training through iterative gradient propagation along the residual pathErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICLR 2025
- An EM Framework for Online Incremental Learning of Semantic SegmentationShipeng Yan, Jiale Zhou, Jiangwei Xie, Songyang Zhang et al.ACM MM 2021 · 32 citations
- CNM-UNet: Continuous Ordinary Differential Equations for Medical Image SegmentationTianqi Xu, Yashi Zhu, Quansong He, Yue Cao et al.AAAI 2026
- Beyond Skip Connection: Pooling and Unpooling Design for Elimination SingularitiesChengkun Sun, Jinqian Pan, Zhuoli Jin, Russell Stevens Terry et al.AAAI 2025 · 1 citation
