Low-Budget Active Learning via Wasserstein Distance: An Integer Programming Approach
Rafid Mahmood, Sanja Fidler, Marc T. Law
摘要
Active learning is the process of training a model with limited labeled data by selecting a core subset of an unlabeled data pool to label. The large scale of data sets used in deep learning forces most sample selection strategies to employ efficient heuristics. This paper introduces an integer optimization problem for selecting a core set that minimizes the discrete Wasserstein distance from the unlabeled pool. We demonstrate that this problem can be tractably solved with a Generalized Benders Decomposition algorithm. Our strategy uses high-quality latent features that can be obtained by unsupervised learning on the unlabeled pool. Numerical results on several data sets show that our optimization approach is competitive with baselines and particularly outperforms them in the low budget regime where less than one percent of the data set is labeled.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 被引用 163 次
- Active Learning Through a Covering LensOfer Yehuda, Avihu Dekel, Guy Hacohen, Daphna WeinshallNeurIPS 2022 · 被引用 102 次
- Optimizing Data Collection for Machine LearningRafid Mahmood, James Lucas, José M. Álvarez, Sanja Fidler 等NeurIPS 2022 · 被引用 39 次
- Knowledge-Aware Federated Active Learning with Non-IID DataYu-Tong Cao, Ye Shi, Baosheng Yu, Jingya Wang 等ICCV 2023 · 被引用 30 次
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and BudgetGuy Hacohen, Daphna WeinshallNeurIPS 2023 · 被引用 23 次
它引用的顶会 Paper3
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 被引用 662 次
相关 Paper
- A Lagrangian Duality Approach to Active LearningJuan Elenter, Navid NaderiAlizadeh, Alejandro RibeiroNeurIPS 2022 · 被引用 31 次
- GALAXY: Graph-based Active Learning at the ExtremeJifan Zhang, Julian Katz-Samuels, Robert D. NowakICML 2022 · 被引用 47 次
- Unsupervised Active Learning via Subspace LearningChangsheng Li, Kaihang Mao, Lingyan Liang, Dongchun Ren 等AAAI 2021 · 被引用 18 次
- Enhancing Semi-Supervised Learning via Representative and Diverse Sample SelectionQian Shao, Jiangrui Kang, Qiyuan Chen, Zepeng Li 等NeurIPS 2024 · 被引用 3 次
- Stochastic Encodings for Active Feature AcquisitionAlexander Luke Ian Norcliffe, Changhee Lee, Fergus Imrie, Mihaela van der Schaar 等ICML 2025
