Cogradient Descent for Bilinear Optimization
Li'an Zhuo, Baochang Zhang, Linlin Yang, Hanlin Chen, Qixiang Ye, David S. Doermann, Rongrong Ji, Guodong Guo
Abstract
Conventional learning methods simplify the bilinear model by regarding two intrinsically coupled factors independently, which degrades the optimization procedure. One reason lies in the insufficient training due to the asynchronous gradient descent, which results in vanishing gradients for the coupled variables. In this paper, we introduce a Cogradient Descent algorithm (CoGD) to address the bilinear problem, based on a theoretical framework to coordinate the gradient of hidden variables via a projection function. We solve one variable by considering its coupling relationship with the other, leading to a synchronous gradient descent to facilitate the optimization procedure. Our algorithm is applied to solve problems with one variable under the sparsity constraint, which is widely used in the learning paradigm. We validate our CoGD considering an extensive set of applications including image reconstruction, inpainting, and network pruning. Experiments show that it improves the state-of-the-art by a significant margin 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 270c6e2a-9531-47d5-9e9e-d8cb00e860b9Cited by top-tier papers5
- SCOP: Scientific Control for Reliable Neural Network PruningYehui Tang, Yunhe Wang, Yixing Xu, Dacheng Tao et al.NeurIPS 2020 · 208 citations
- Searching for Low-Bit Weights in Quantized Neural NetworksZhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu et al.NeurIPS 2020 · 103 citations
- GradAug: A New Regularization Method for Deep Neural NetworksTaojiannan Yang, Sijie Zhu, Chen ChenNeurIPS 2020 · 43 citations
- IDARTS: Interactive Differentiable Architecture SearchSong Xue, Runqi Wang, Baochang Zhang, Tian Wang et al.ICCV 2021 · 10 citations
- Layer-Wise Searching for 1-Bit DetectorsSheng Xu, Junhe Zhao, Jinhu Lü, Baochang Zhang et al.CVPR 2021
Builds on3
- Noise Flow: Noise Modeling With Conditional Normalizing FlowsAbdelrahman Abdelhamed, Marcus A. Brubaker, Michael S. BrownICCV 2019 · 199 citations
- Accelerate CNN via Recursive Bayesian PruningYuefu Zhou, Ya Zhang, Yanfeng Wang, Qi TianICCV 2019 · 64 citations
- Solving Vision Problems via FilteringSean I. Young, Aous Thabit Naman, Bernd Girod, David TaubmanICCV 2019 · 4 citations
Related papers
- Granger Components Analysis: Unsupervised learning of latent temporal dependenciesJacek DmochowskiNeurIPS 2023 · 2 citations
- GAN-Based Projector for Faster Recovery With Convergence Guarantees in Linear Inverse ProblemsAnkit Raj, Yuqi Li, Yoram BreslerICCV 2019 · 61 citations
- Nesterov Meets Optimism: Rate-Optimal Separable Minimax OptimizationChris Junchi Li, Huizhuo Yuan, Gauthier Gidel, Quanquan Gu et al.ICML 2023 · 8 citations
- GDA-AM: On the Effectiveness of Solving Min-Imax Optimization via Anderson MixingHuan He, Shifan Zhao, Yuanzhe Xi, Joyce C. Ho et al.ICLR 2022 · 12 citations
- Smoothing Proximal Gradient Methods for Nonsmooth Sparsity Constrained Optimization: Optimality Conditions and Global ConvergenceGanzhao YuanICML 2024 · 6 citations
