Learning Proximal Operators to Discover Multiple Optima
Lingxiao Li, Noam Aigerman, Vladimir G. Kim, Jiajin Li, Kristjan H. Greenewald, Mikhail Yurochkin, Justin Solomon
Abstract
Finding multiple solutions of non-convex optimization problems is a ubiquitous yet challenging task. Most past algorithms either apply single-solution optimization methods from multiple random initial guesses or search in the vicinity of found solutions using ad hoc heuristics. We present an end-to-end method to learn the proximal operator of a family of training problems so that multiple local minima can be quickly obtained from initial guesses by iterating the learned operator, emulating the proximal-point algorithm that has fast convergence. The learned proximal operator can be further generalized to recover multiple optima for unseen problems at test time, enabling applications such as object detection. The key ingredient in our formulation is a proximal regularization term, which elevates the convexity of our training loss: by applying recent theoretical results, we show that for weakly-convex objectives with Lipschitz gradients, training of the proximal operator converges globally with a practical degree of over-parameterization. We further present an exhaustive benchmark for multi-solution optimization to demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2aa2bf8e-aba4-4ac4-9842-e582f9b5db86Cited by top-tier papers2
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin et al.ICML 2023 · 18 citations
- Beyond Scores: Proximal Diffusion ModelsZhenghan Fang, Mateo Díaz, Sam Buchanan, Jeremias SulamNeurIPS 2025 · 6 citations
Builds on6
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Large-Scale Wasserstein Gradient FlowsPetr Mokrov, Alexander Korotin, Lingxiao Li, Aude Genevay et al.NeurIPS 2021 · 112 citations
- CvxNet: Learnable Convex DecompositionBoyang Deng, Kyle Genova, Soroosh Yazdani, Sofien Bouaziz et al.CVPR 2020
- NeRD: Neural 3D Reflection Symmetry DetectorYichao Zhou, Shichen Liu, Yi MaCVPR 2021
Related papers
- What's in a Prior? Learned Proximal Networks for Inverse ProblemsZhenghan Fang, Sam Buchanan, Jeremias SulamICLR 2024 · 27 citations
- Learn2Hop: Learned Optimization on Rough LandscapesAmil Merchant, Luke Metz, Samuel S. Schoenholz, Ekin D. CubukICML 2021 · 19 citations
- Effective Meta-Regularization by Kernelized Proximal RegularizationWeisen Jiang, James T. Kwok, Yu ZhangNeurIPS 2021 · 9 citations
- Newton Informed Neural Operator for Solving Nonlinear Partial Differential EquationsWenrui Hao, Xinliang Liu, Yahong YangNeurIPS 2024 · 22 citations
- How to Fill the Optimum Set? Population Gradient Descent with Harmless DiversityChengyue Gong, Lemeng Wu, Qiang LiuICML 2022 · 4 citations
