Automatic and Harmless Regularization with Constrained and Lexicographic Optimization: A Dynamic Barrier Approach
Chengyue Gong, Xingchao Liu, Qiang Liu
摘要
Many machine learning tasks have to make a trade-off between two loss functions, typically the main data-fitness loss and an auxiliary loss. The most widely used approach is to optimize the linear combination of the objectives, which, however, requires manual tuning of the combination coefficient and is theoretically unsuitable for non-convex functions. In this work, we consider constrained optimization as a more principled approach for trading off two losses, with a special emphasis on lexicographic (lexico) optimization, a degenerated limit of constrained optimization which optimizes a secondary loss inside the optimal set of the main loss. We propose a dynamic barrier gradient descent algorithm which provides a unified solution of both constrained and lexicographic optimization. We establish the convergence of the method for general non-convex functions. Through a number of experiments on real-world deep learning tasks, we show that 1) lexico optimization provides a tuning-free approach to incorporating side loss functions without hurting the main objective, and 2) constrained and lexico optimization combined provide an automatic approach to profiling Pareto sets, especially in non-convex problems on which linear combination methods fail.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- BOME! Bilevel Optimization Made Easy: A Simple First-Order ApproachBo Liu, Mao Ye, Stephen Wright, Peter Stone 等NeurIPS 2022 · 被引用 170 次
- A Fully First-Order Method for Stochastic Bilevel OptimizationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICML 2023 · 被引用 123 次
- On Penalty-based Bilevel Gradient Descent MethodHan Shen, Tianyi ChenICML 2023 · 被引用 105 次
- A Multi-objective / Multi-task Learning Framework Induced by Pareto StationarityMichinari Momma, Chaosheng Dong, Jia LiuICML 2022 · 被引用 62 次
- Averaged Method of Multipliers for Bi-Level Optimization without Lower-Level Strong ConvexityRisheng Liu, Yaohua Liu, Wei Yao, Shangzhi Zeng 等ICML 2023 · 被引用 37 次
它引用的顶会 Paper8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor 等NeurIPS 2020 · 被引用 506 次
- Minimax Pareto Fairness: A Multi Objective PerspectiveNatalia Martínez, Martín Bertrán, Guillermo SapiroICML 2020 · 被引用 232 次
相关 Paper
- Multi-Task Learning with User Preferences: Gradient Descent with Controlled Ascent in Pareto OptimizationDebabrata Mahapatra, Vaibhav RajanICML 2020 · 被引用 182 次
- Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference GuidanceLisha Chen, Quan Xiao, Ellen Hidemi Fukuda, Xinyi Chen 等ICML 2025
- Optimizing Neural Networks with Gradient Lexicase SelectionLi Ding, Lee SpectorICLR 2022 · 被引用 21 次
- A Single-Loop Gradient Descent and Perturbed Ascent Algorithm for Nonconvex Functional Constrained OptimizationSongtao LuICML 2022 · 被引用 27 次
- Learning with Privileged TasksYuru Song, Zan Lou, Shan You, Erkun Yang 等ICCV 2021 · 被引用 3 次
