From Selection to Generation: A Survey of LLM-based Active Learning
Yu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen, Franck Dernoncourt, Branislav Kveton, Tong Yu
Abstract
Active Learning (AL) has been a powerful paradigm for improving model efficiency and performance by selecting the most informative data points for labeling and training. In recent active learning frameworks, Large Language Models (LLMs) have been employed not only for selection but also for generating entirely new data instances and providing more cost-effective annotations. Motivated by the increasing importance of high-quality data and efficient model training in the era of LLMs, we present a comprehensive survey on LLM-based Active Learning. We introduce an intuitive taxonomy that categorizes these techniques and discuss the transformative roles LLMs can play in the active learning loop. We further examine the impact of AL on LLM learning paradigms and its applications across various domains. Finally, we identify open challenges and propose future research directions. This survey aims to serve as an up-to-date resource for researchers and practitioners seeking to gain an intuitive understanding of LLM-based AL techniques and deploy them to new applications. * Equal contributions Model Training LLM-based Selection / Generation selected new generated 𝑥 ! data point data point LLM-based Annotation or Human Annotation initial data u n la b e le d d a ta 𝑥 " selected/gen. data points (𝑥 " , 𝑦 " ) (𝑥 ! , 𝑦 ! ) Class General Mechanism Description Querying (Section 3) Traditional Selection (Sec. 3.1) This class of techniques uses traditional selection such as uncertainty sampling, disagreement, gradient-based sampling, and so on. LLM-based Selection (Sec. 3.2) The class of LLM-based selection techniques focus on using LLMs to select the instances. LLM-based Generation (Sec. 3.3) The class of LLM-based generation techniques focus on generating novel instances. Hybrid (Sec. 3.4) Combines advantages of both LLM-based selection and generation Annotation (Section 4) Human Annotation (Sec. 4.1) Traditional human annotation simply refers to using humans to annotate the selected or generated instances, which is costly. LLM-based Annotation (Sec. 4.2) The class of LLM-based annotation techniques focus on leveraging LLMs for annotation and evaluation. This class of techniques are far cheaper than human annotation. Hybrid (Sec. 4.3) This class of techniques aim to leverage the advantages of both humans and LLMs for optimal annotations while minimizing cost
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf163210-89a1-4132-afe9-9bc983e78382Cited by top-tier papers10
- MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM SafetyJialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen et al.ICML 2026 · 5 citations
- ProbeLLM: Automating Principled Diagnosis of LLM FailuresYue Huang, Zhengzhe Jiang, Yuchen Ma, Yu Jiang et al.ICML 2026 · 4 citations
- Principle-Evolvable Scientific Discovery via Uncertainty MinimizationYingming Pu, Tao LIN, Hongyu ChenICML 2026 · 4 citations
- Uncertainty-Aware Clarification in LLM Agents with Information GainMengyi DENG, Zhiwei Li, Xin Li, Tingyu ZHU et al.ICML 2026 · 1 citation
- SAND: Boosting LLM Agents with Self-Taught Action DeliberationYu Xia, Yiran Shen, Junda Wu, Tong Yu et al.EMNLP 2025 · 1 citation
Builds on20
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Batch Active Learning at ScaleGui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas et al.NeurIPS 2021 · 220 citations
- Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference AdjustmentRui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu et al.ICML 2024 · 144 citations
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra et al.CHI 2024 · 127 citations
- Gone Fishing: Neural Active Learning with Fisher EmbeddingsJordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, Sham M. KakadeNeurIPS 2021 · 124 citations
Related papers
- Next Generation Active Learning: Mixture of LLMs in the LoopYuanyuan Qi, Xiaohao Yang, Jueqing Lu, Guoxiang Guo et al.AAAI 2026
- Active Learning for Natural Language GenerationYotam Perlitz, Ariel Gera, Michal Shmueli-Scheuer, Dafna Sheinwald et al.EMNLP 2023 · 2 citations
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- A Survey on Efficient Large Language Model Training: From Data-centric PerspectivesJunyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao et al.ACL 2025 · 12 citations
- LLMs Are Noisy Oracles! LLM-based Noise-aware Graph Active Learning for Node ClassificationZeang Sheng, Weiyang Guo, Yingxia Shao, Wentao Zhang et al.KDD 2025 · 1 citation
