Preference-Guided Reflective Sampling for Aligning Language Models
Hai Ye, Hwee Tou Ng
Abstract
Iterative data generation and model re-training can effectively align large language models (LLMs) to human preferences. The process of data sampling is crucial, as it significantly influences the success of policy improvement. Repeated random sampling is a widely used method that independently queries the model multiple times to generate outputs. In this work, we propose a more effective sampling method, named Preference-Guided Reflective Sampling (PRS). Unlike random sampling, PRS employs a tree-based generation framework to enable more efficient sampling. It leverages adaptive self-refinement techniques to better explore the sampling space. By specifying user preferences in natural language, PRS can further optimize response generation according to these preferences. As a result, PRS can align models to diverse user preferences. Our experiments demonstrate that PRS generates higher-quality responses with significantly higher rewards. On AlpacaEval and Arena-Hard, PRS substantially outperforms repeated random sampling in bestof-N sampling. Moreover, PRS shows strong performance when applied in iterative offline RL training 1 . * Provide references or sources to support each claim made in the response ... * Break down the response into smaller, more manageable sections ... 1. Timeline chart: This is a graphical representation of events or milestones in chronological order. ... 2. Gantt chart: A Gantt chart is a type of bar chart used to show a schedule of a ... References * [1]: Wikipedia, "Timeline," <https ://en.wikipedia.org/wiki/Timeline> ... Response Feedback Refined Response 2. Sample response 3. Provide feedback 4. Revise response 1. Add explicit preference (a) Reflective Refinement (b) Tree-based Generation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Finding the Sweet Spot: Preference Data Construction for Scaling Preference OptimizationYao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng et al.ACL 2025 · 8 citations
- Think&Cite: Improving Attributed Text Generation with Self-Guided Tree Search and Progress Reward ModelingJunyi Li, Hwee Tou NgACL 2025 · 5 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- HPS: Hard Preference Sampling for Human Preference AlignmentXiandong Zou, Wanyu Lin, Yuchen Li, Pan ZhouICML 2025
- Adaptive Batch-Wise Sample Scheduling for Direct Preference OptimizationZixuan Huang, Yikun Ban, Lean Fu, Xiaojie Li et al.NeurIPS 2025 · 14 citations
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu et al.AAAI 2024 · 357 citations
- Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference ModelJunshu Pan, Wei Shen, Shulin Huang, Qiji Zhou et al.AAAI 2026 · 7 citations
- Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay PerspectiveRuichen Shao, Bei Li, Gangao Liu, Yang Chen et al.ICLR 2025
