A Lagrangian Duality Approach to Active Learning
Juan Elenter, Navid NaderiAlizadeh, Alejandro Ribeiro
Abstract
We consider the pool-based active learning problem, where only a subset of the training data is labeled, and the goal is to query a batch of unlabeled samples to be labeled so as to maximally improve model performance. We formulate the problem using constrained learning, where a set of constraints bounds the performance of the model on labeled samples. Considering a primal-dual approach, we optimize the primal variables, corresponding to the model parameters, as well as the dual variables, corresponding to the constraints. As each dual variable indicates how significantly the perturbation of the respective constraint affects the optimal value of the objective function, we use it as a proxy of the informativeness of the corresponding training sample. Our approach, which we refer to as Active Learning via Lagrangian dualitY, or ALLY, leverages this fact to select a diverse set of unlabeled samples with the highest estimated dual variables as our query set. We demonstrate the benefits of our approach in a variety of classification and regression tasks and discuss its limitations depending on the capacity of the model used and the degree of redundancy in the dataset. We also examine the impact of the distribution shift induced by active sampling and show that ALLY can be used in a generative mode to create novel, maximally-informative samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Algorithm Selection for Deep Active Learning with Imbalanced DatasetsJifan Zhang, Shuai Shao, Saurabh Verma, Robert D. NowakNeurIPS 2023 · 38 citations
- Resilient Constrained LearningIgnacio Hounie, Alejandro Ribeiro, Luiz F. O. ChamonNeurIPS 2023 · 20 citations
- Automatic Data Augmentation via Invariance-Constrained LearningIgnacio Hounie, Luiz F. O. Chamon, Alejandro RibeiroICML 2023 · 20 citations
- Near-Optimal Solutions of Constrained Learning ProblemsJuan Elenter, Luiz F. O. Chamon, Alejandro RibeiroICLR 2024 · 10 citations
- Balancing Act: Constraining Disparate Impact in Sparse ModelsMeraj Hashemizadeh, Juan Ramirez, Rohan Sukumaran, Golnoosh Farnadi et al.ICLR 2024 · 9 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
Related papers
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and BudgetGuy Hacohen, Daphna WeinshallNeurIPS 2023 · 23 citations
- Influence Selection for Active LearningZhuoming Liu, Hao Ding, Huaping Zhong, Weijia Li et al.ICCV 2021 · 125 citations
- Discover-Then-Rank Unlabeled Support Vectors in the Dual Space for Multi-Class Active LearningDayou Yu, Weishi Shi, Qi YuICML 2023 · 1 citation
- Provably Neural Active Learning Succeeds via Prioritizing Perplexing SamplesDake Bu, Wei Huang, Taiji Suzuki, Ji Cheng et al.ICML 2024 · 5 citations
- Low-Budget Active Learning via Wasserstein Distance: An Integer Programming ApproachRafid Mahmood, Sanja Fidler, Marc T. LawICLR 2022 · 44 citations
