Self-Boosting Large Language Models with Synthetic Preference Data
Qingxiu Dong, Li Dong, Xingxing Zhang, Zhifang Sui, Furu Wei
Abstract
Through alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb8f75da-d41d-49a8-8f3f-5ae6fcfeadbdCited by top-tier papers12
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-TrainingLai Wei, Yuting Li, Chen Wang, Yue Wang et al.NeurIPS 2025 · 28 citations
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen et al.NeurIPS 2025 · 13 citations
- A Survey on Efficient Large Language Model Training: From Data-centric PerspectivesJunyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao et al.ACL 2025 · 12 citations
- Finding the Sweet Spot: Preference Data Construction for Scaling Preference OptimizationYao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng et al.ACL 2025 · 8 citations
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
Related papers
- Self-Play Preference Optimization for Language Model AlignmentYue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji et al.ICLR 2025
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM AlignmentXiaoyang Cao, Zelai Xu, Mo Guang, Kaiwen Long et al.ICLR 2026 · 4 citations
- Latent Preference Coding: Aligning Large Language Models via Discrete Latent CodesZhuocheng Gong, Jian Guan, Wei Wu, Huishuai Zhang et al.ICML 2025
- Improving Model Alignment Through Collective Intelligence of Open-Source ModelsJunlin Wang, Roy Xie, Shang Zhu, Jue Wang et al.ICML 2025
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen et al.ACL 2026 · 2 citations
