ALVIN: Active Learning Via INterpolation
Michalis Korakakis, Andreas Vlachos, Adrian Weller
Abstract
Active Learning aims to minimize annotation effort by selecting the most useful instances from a pool of unlabeled data. However, typical active learning methods overlook the presence of distinct example groups within a class, whose prevalence may vary, e.g., in occupation classification datasets certain demographics are disproportionately represented in specific classes. This oversight causes models to rely on shortcuts for predictions, i.e., spurious correlations between input attributes and labels occurring in well-represented groups. To address this issue, we propose Active Learning Via INterpolation (ALVIN), which conducts intra-class interpolations between examples from underrepresented and well-represented groups to create anchors, i.e., artificial points situated between the example groups in the representation space. By selecting instances close to the anchors for annotation, ALVIN identifies informative examples exposing the model to regions of the representation space that counteract the influence of shortcuts. Crucially, since the model considers these examples to be of high certainty, they are likely to be ignored by typical active learning methods. Experimental results on six datasets encompassing sentiment analysis, natural language inference, and paraphrase detection demonstrate that ALVIN outperforms state-of-the-art active learning methods in both in-distribution and out-of-distribution generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aaddbf5d-6ff2-414b-bd0b-65c17e1a84ccCited by top-tier papers2
- Unsupervised Process-Aware Coreset Selection for In-Context LearningWei Zheng, Zijie Wang, Xin Li, Bin Gong et al.ICML 2026
- Mitigating Shortcut Learning with InterpoLated LearningMichalis Korakakis, Andreas Vlachos, Adrian WellerACL 2025
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
Related papers
- GALAXY: Graph-based Active Learning at the ExtremeJifan Zhang, Julian Katz-Samuels, Robert D. NowakICML 2022 · 47 citations
- Contrastive Coding for Active Learning under Class Distribution MismatchPan Du, Suyun Zhao, Hui Chen, Shuwen Chai et al.ICCV 2021 · 50 citations
- Counterfactual Active Learning for Out-of-Distribution GeneralizationXun Deng, Wenjie Wang, Fuli Feng, Hanwang Zhang et al.ACL 2023 · 10 citations
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch et al.EMNLP 2020 · 144 citations
- Class Balance Matters to Active Class-Incremental LearningZitong Huang, Ze Chen, Yuanze Li, Bowen Dong et al.ACM MM 2024 · 8 citations
