Self-Boosting Large Language Models with Synthetic Preference Data
Qingxiu Dong, Li Dong, Xingxing Zhang, Zhifang Sui, Furu Wei
摘要
Through alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 被引用 59 次
- First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-TrainingLai Wei, Yuting Li, Chen Wang, Yue Wang 等NeurIPS 2025 · 被引用 28 次
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen 等NeurIPS 2025 · 被引用 13 次
- A Survey on Efficient Large Language Model Training: From Data-centric PerspectivesJunyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao 等ACL 2025 · 被引用 12 次
- Finding the Sweet Spot: Preference Data Construction for Scaling Preference OptimizationYao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng 等ACL 2025 · 被引用 8 次
它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- Self-Play Preference Optimization for Language Model AlignmentYue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji 等ICLR 2025
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM AlignmentXiaoyang Cao, Zelai Xu, Mo Guang, Kaiwen Long 等ICLR 2026 · 被引用 4 次
- Latent Preference Coding: Aligning Large Language Models via Discrete Latent CodesZhuocheng Gong, Jian Guan, Wei Wu, Huishuai Zhang 等ICML 2025
- Improving Model Alignment Through Collective Intelligence of Open-Source ModelsJunlin Wang, Roy Xie, Shang Zhu, Jue Wang 等ICML 2025
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen 等ACL 2026 · 被引用 2 次
