A Neural Pre-Conditioning Active Learning Algorithm to Reduce Label Complexity
Seo Taek Kong, Soomin Jeon, Dongbin Na, Jaewon Lee, Hong-Seok Lee, Kyu-Hwan Jung
Abstract
Deep learning (DL) algorithms rely on massive amounts of labeled data. Semi-supervised learning (SSL) and active learning (AL) aim to reduce this label complexity by leveraging unlabeled data or carefully acquiring labels, respectively. In this work, we primarily focus on designing an AL algorithm but first argue for a change in how AL algorithms should be evaluated. Although unlabeled data is readily available in pool-based AL, AL algorithms are usually evaluated by measuring the increase in supervised learning (SL) performance at consecutive acquisition steps. Because this measures performance gains from both newly acquired instances and newly acquired labels, we propose to instead evaluate the label efficiency of AL algorithms by measuring the increase in SSL performance at consecutive acquisition steps. After surveying tools that can be used to this end, we propose our neural pre-conditioning (NPC) algorithm inspired by a Neural Tangent Kernel (NTK) analysis. Our algorithm incorporates the classifier's uncertainty on unlabeled data and penalizes redundant samples within candidate batches to efficiently acquire a diverse set of informative labels. Furthermore, we prove that NPC improves downstream training in the large-width regime in a manner previously observed to correlate with generalization. Comparisons with other AL algorithms show that a state-of-the-art SSL algorithm coupled with NPC can achieve high performance using very few labeled data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 875f4e14-cfc4-4d36-af1a-5a7ce15b6ddaCited by top-tier papers2
- BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad RangesHoyong Choi, Nohyun Ki, Hye Won ChungICML 2024 · 9 citations
- Continuous Learning for Android Malware DetectionYizheng Chen, Zhoujie Ding, David A. WagnerUSENIX Security 2023
Builds on7
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin et al.ICLR 2020 · 469 citations
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised LearningJaehyung Kim, Youngbum Hur, Sejun Park, Eunho Yang et al.NeurIPS 2020 · 209 citations
Related papers
- Making Look-Ahead Active Learning Strategies Feasible with Neural Tangent KernelsMohamad Amin Mohamadi, Wonho Bae, Danica J. SutherlandNeurIPS 2022 · 32 citations
- Deep Active Learning for Biased Datasets via Fisher Kernel Self-SupervisionDenis A. Gudovskiy, Alec Hodgkinson, Takuya Yamaguchi, Sotaro TsukizawaCVPR 2020
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- Navigating the Pitfalls of Active Learning Evaluation: A Systematic Framework for Meaningful Performance AssessmentCarsten T. Lüth, Till J. Bungert, Lukas Klein, Paul F. JaegerNeurIPS 2023 · 32 citations
- State-Relabeling Adversarial Active LearningBeichen Zhang, Liang Li, Shijie Yang, Shuhui Wang et al.CVPR 2020
