Margin-based sampling in high dimensions: When being active is less efficient than staying passive
Alexandru Tifrea, Jacob Clarysse, Fanny Yang
Abstract
It is widely believed that given the same labeling budget, active learning (AL) algorithms like margin-based active learning achieve better predictive performance than passive learning (PL), albeit at a higher computational cost. Recent empirical evidence suggests that this added cost might be in vain, as margin-based AL can sometimes perform even worse than PL. While existing works offer different explanations in the low-dimensional regime, this paper shows that the underlying mechanism is entirely different in high dimensions: we prove for logistic regression that PL outperforms margin-based AL even for noiseless data and when using the Bayes optimal decision boundary for sampling. Insights from our proof indicate that this high-dimensional phenomenon is exacerbated when the separation between the classes is small. We corroborate this intuition with experiments on 20 high-dimensional datasets spanning a diverse range of applications, from finance and histology to chemistry and computer vision. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Active Test-time Vision-Language NavigationHeeju Ko, Sung June Kim, Gyeongrok Oh, Jeongyoon Yoon et al.NeurIPS 2025 · 10 citations
- MER-Inspector: Assessing Model Extraction Risks from An Attack-Agnostic PerspectiveXinwei Zhang, Haibo Hu, Qingqing Ye, Li Bai et al.WWW 2025 · 5 citations
Builds on4
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 163 citations
- On Statistical Bias In Active Learning: How and When to Fix ItSebastian Farquhar, Yarin Gal, Tom RainforthICLR 2021 · 96 citations
- Interpolation can hurt robust generalization even when there is no noiseKonstantin Donhauser, Alexandru Tifrea, Michael Aerni, Reinhard Heckel et al.NeurIPS 2021 · 18 citations
Related papers
- Constants Matter: The Performance Gains of Active LearningStephen O. Mussmann, Sanjoy DasguptaICML 2022 · 1 citation
- FIRAL: An Active Learning Algorithm for Multinomial Logistic RegressionYouguang Chen, George BirosNeurIPS 2023 · 3 citations
- Online Active Learning with Surrogate Loss FunctionsGiulia DeSalvo, Claudio Gentile, Tobias Sommer ThuneNeurIPS 2021 · 9 citations
- Improved Algorithm for Deep Active Learning under Imbalance via Optimal SeparationShyam Nuggehalli, Jifan Zhang, Lalit K. Jain, Robert D. NowakICML 2025
- Improved Algorithms for Neural Active LearningYikun Ban, Yuheng Zhang, Hanghang Tong, Arindam Banerjee et al.NeurIPS 2022 · 18 citations
