Margin-based sampling in high dimensions: When being active is less efficient than staying passive
Alexandru Tifrea, Jacob Clarysse, Fanny Yang
摘要
It is widely believed that given the same labeling budget, active learning (AL) algorithms like margin-based active learning achieve better predictive performance than passive learning (PL), albeit at a higher computational cost. Recent empirical evidence suggests that this added cost might be in vain, as margin-based AL can sometimes perform even worse than PL. While existing works offer different explanations in the low-dimensional regime, this paper shows that the underlying mechanism is entirely different in high dimensions: we prove for logistic regression that PL outperforms margin-based AL even for noiseless data and when using the Bayes optimal decision boundary for sampling. Insights from our proof indicate that this high-dimensional phenomenon is exacerbated when the separation between the classes is small. We corroborate this intuition with experiments on 20 high-dimensional datasets spanning a diverse range of applications, from finance and histology to chemistry and computer vision. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Active Test-time Vision-Language NavigationHeeju Ko, Sung June Kim, Gyeongrok Oh, Jeongyoon Yoon 等NeurIPS 2025 · 被引用 10 次
- MER-Inspector: Assessing Model Extraction Risks from An Attack-Agnostic PerspectiveXinwei Zhang, Haibo Hu, Qingqing Ye, Li Bai 等WWW 2025 · 被引用 5 次
它引用的顶会 Paper4
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 被引用 163 次
- On Statistical Bias In Active Learning: How and When to Fix ItSebastian Farquhar, Yarin Gal, Tom RainforthICLR 2021 · 被引用 96 次
- Interpolation can hurt robust generalization even when there is no noiseKonstantin Donhauser, Alexandru Tifrea, Michael Aerni, Reinhard Heckel 等NeurIPS 2021 · 被引用 18 次
相关 Paper
- Constants Matter: The Performance Gains of Active LearningStephen O. Mussmann, Sanjoy DasguptaICML 2022 · 被引用 1 次
- FIRAL: An Active Learning Algorithm for Multinomial Logistic RegressionYouguang Chen, George BirosNeurIPS 2023 · 被引用 3 次
- Online Active Learning with Surrogate Loss FunctionsGiulia DeSalvo, Claudio Gentile, Tobias Sommer ThuneNeurIPS 2021 · 被引用 9 次
- Improved Algorithm for Deep Active Learning under Imbalance via Optimal SeparationShyam Nuggehalli, Jifan Zhang, Lalit K. Jain, Robert D. NowakICML 2025
- Improved Algorithms for Neural Active LearningYikun Ban, Yuheng Zhang, Hanghang Tong, Arindam Banerjee 等NeurIPS 2022 · 被引用 18 次
