Deep Bayesian Active Learning for Preference Modeling in Large Language Models
Luckeciano Carvalho Melo, Panagiotis Tigas, Alessandro Abate, Yarin Gal
摘要
Leveraging human preferences for steering the behavior of Large Language Models (LLMs) has demonstrated notable success in recent years. Nonetheless, data selection and labeling are still a bottleneck for these systems, particularly at large scale. Hence, selecting the most informative points for acquiring human feedback may considerably reduce the cost of preference labeling and unleash the further development of LLMs. Bayesian Active Learning provides a principled framework for addressing this challenge and has demonstrated remarkable success in diverse settings. However, previous attempts to employ it for Preference Modeling did not meet such expectations. In this work, we identify that naive epistemic uncertainty estimation leads to the acquisition of redundant samples. We address this by proposing the Bayesian Active Learner for Preference Modeling (BAL-PM), a novel stochastic acquisition policy that not only targets points of high epistemic uncertainty according to the preference model but also seeks to maximize the entropy of the acquired prompt distribution in the feature space spanned by the employed LLM. Notably, our experiments demonstrate that BAL-PM requires 33% to 68% fewer preference labels in two popular human preference datasets and exceeds previous stochastic Bayesian acquisition policies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- From Selection to Generation: A Survey of LLM-based Active LearningYu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu 等ACL 2025 · 被引用 18 次
- CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk PredictionSizhe Wang, Ziqi Xu, Claire Najjuuko, Charles Alba 等ACL 2026
- ActiveUltraFeedback: Efficient Preference Data Generation using Active LearningDavit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich 等ICML 2026
- Dimension-Aware Active Annotation for Aesthetic Perception via Multi-Agent Human-AI CollaborationYe Zhang, Jinlong He, Dongjie Wang, Yupeng Zhou 等AAAI 2026
- A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and SolutionsZhiyin Yu, Yuchen Mou, Juncheng Yan, Junyu Luo 等ACL 2026
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
相关 Paper
- Active Preference Learning for Large Language ModelsWilliam Muldrew, Peter Hayes, Mingtian Zhang, David BarberICML 2024 · 被引用 53 次
- Comparison-based Active Preference Learning for Multi-dimensional PersonalizationMinhyeon Oh, Seungjoon Lee, Jungseul OkACL 2025 · 被引用 1 次
- LILO: Bayesian Optimization with Natural Language FeedbackKatarzyna Kobalczyk, Zhiyuan Lin, Benjamin Letham, Zhuokai Zhao 等ICML 2026 · 被引用 2 次
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 被引用 3 次
- Epistemic Uncertainty Estimation in Regression Ensemble Models with Pairwise Epistemic EstimatorsLucas Berry, David MegerNeurIPS 2025 · 被引用 6 次
