When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
Amirabbas Afzali, Myeongho Jeon, Maria Brbic
摘要
Preference alignment is an essential step in adapting large language models (LLMs) to human values, but existing approaches typically depend on costly human annotations or large-scale API-based models. We explore whether a weak LLM can instead act as an effective annotator. We surprisingly find that selecting only a subset of a weak LLM's highly confident samples leads to substantially better performance than using full human annotations. Building on this insight, we propose Confidence-Weighted Preference Optimization (CW-PO), a general framework that re-weights training samples by a weak LLM’s confidence and can be applied across different preference optimization objectives. Notably, the model aligned by CW-PO with just 20% of human annotations outperforms the model trained with 100% of annotations under standard DPO. These results suggest that weak LLMs, when paired with confidence weighting, can dramatically reduce the cost of preference alignment while even outperforming methods trained on fully human-labeled data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
相关 Paper
- Your Weak LLM is Secretly a Strong Teacher for AlignmentLeitian Tao, Yixuan LiICLR 2025
- Preference-Strength-Aware Self-Improving Alignment with Generative Preference ModelsYuanzhao Zhai, Zhuo Zhang, Cheng Yang, Kele Xu 等SIGIR 2025
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient AlignmentXiaoqiang Lin, Arun Verma, Zhongxiang Dai, Daniela Rus 等ICLR 2026 · 被引用 12 次
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference OptimizationJunming Yang, Ning Xu, Biao Liu, Shiqi Qiao 等ICLR 2026 · 被引用 3 次
- Reliability-Aware LLM Alignment from Inconsistent Human FeedbackJingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma 等ICML 2026
