Optimizing Nondecomposable Data Dependent Regularizers via Lagrangian Reparameterization Offers Significant Performance and Efficiency Gains
Sathya N. Ravi, Abhay Venkatesh, Glenn Moo Fung, Vikas Singh
Abstract
Data dependent regularization is known to benefit a wide variety of problems in machine learning. Often, these regularizers cannot be easily decomposed into a sum over a finite number of terms, e.g., a sum over individual example-wise terms. The F β measure, Area under the ROC curve (AUCROC) and Precision at a fixed recall (P@R) are some prominent examples that are used in many applications. We find that for most medium to large sized datasets, scalability issues severely limit our ability in leveraging the benefits of such regularizers. Importantly, the key technical impediment despite some recent progress is that, such objectives remain difficult to optimize via backpropapagation procedures. While an efficient general-purpose strategy for this problem still remains elusive, in this paper, we show that for many data-dependent nondecomposable regularizers that are relevant in applications, sizable gains in efficiency are possible with minimal code-level changes; in other words, no specialized tools or numerical schemes are needed. Our procedure involves a reparameterization followed by a partial dualization - this leads to a formulation that has provably cheap projection operators. We present a detailed analysis of runtime and convergence properties of our algorithm. On the experimental side, we show that a direct use of our scheme significantly improves the state of the art IOU measures reported for MSCOCO Stuff segmentation dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Physarum Powered Differentiable Linear Programming Layers and ApplicationsZihang Meng, Sathya N. Ravi, Vikas SinghAAAI 2021 · 5 citations
- Differentiable Optimization of Generalized Nondecomposable Functions using Linear ProgramsZihang Meng, Lopamudra Mukherjee, Yichao Wu, Vikas Singh et al.NeurIPS 2021 · 1 citation
Related papers
- Large-scale Optimization of Partial AUC in a Range of False Positive RatesYao Yao, Qihang Lin, Tianbao YangNeurIPS 2022 · 24 citations
- Finite-Sum Coupled Compositional Stochastic Optimization: Theory and ApplicationsBokun Wang, Tianbao YangICML 2022 · 38 citations
- Quadruply Stochastic Gradient Method for Large Scale Nonlinear Semi-Supervised Ordinal Regression AUC OptimizationWanli Shi, Bin Gu, Xiang Li, Heng HuangAAAI 2020 · 13 citations
- Exploring the Algorithm-Dependent Generalization of AUPRC Optimization with List StabilityPeisong Wen, Qianqian Xu, Zhiyong Yang, Yuan He et al.NeurIPS 2022 · 15 citations
- DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity MeasuresHuanrui Yang, Wei Wen, Hai LiICLR 2020 · 109 citations
