RLVF: Learning from Verbal Feedback without Overgeneralization
Moritz Stephan, Alexander Khazatsky, Eric Mitchell, Annie S. Chen, Sheryl Hsu, Archit Sharma, Chelsea Finn
Abstract
The diversity of contexts in which large language models (LLMs) are deployed requires the ability to modify or customize default model behaviors to incorporate nuanced requirements and preferences. A convenient interface to specify such model adjustments is high-level verbal feedback, such as"Don't use emojis when drafting emails to my boss."However, while writing high-level feedback is far simpler than collecting annotations for reinforcement learning from human feedback (RLHF), we find that simply prompting a model with such feedback leads to overgeneralization of the feedback to contexts where it is not relevant. We study the problem of incorporating verbal feedback without such overgeneralization, inspiring a new method Contextualized Critiques with Constrained Preference Optimization (C3PO). C3PO uses a piece of high-level feedback to generate a small synthetic preference dataset specifying how the feedback should (and should not) be applied. It then fine-tunes the model in accordance with the synthetic preference data while minimizing the divergence from the original model for prompts where the feedback does not apply. Our experimental results indicate that our approach effectively applies verbal feedback to relevant scenarios while preserving existing behaviors for other contexts. For both human- and GPT-4-generated high-level feedback, C3PO effectively adheres to the given feedback comparably to in-context baselines while reducing overgeneralization by 30%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1d02221-c6db-4cbd-8b8f-9c80aa30f99bCited by top-tier papers4
- Characterizing Photorealism and Artifacts in Diffusion Model-Generated ImagesNegar Kamali, Karyn Nakamura, Aakriti Kumar, Angelos Chatzimparmpas et al.CHI 2025 · 23 citations
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann et al.ICML 2026
- Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasNishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng et al.ACL 2025
- Let Me Teach You: Pedagogical Foundations of Feedback for Language ModelsBeatriz Borges, Niket Tandon, Tanja Käser, Antoine BosselutEMNLP 2024
Builds on12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning et al.ICML 2022 · 520 citations
Related papers
- Rule Based Rewards for Language Model SafetyTong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam et al.NeurIPS 2024 · 159 citations
- Explicit Preference Optimization: No Need for an Implicit Reward ModelXiangkun Hu, Lemin Kong, Tong He, David WipfICML 2025
- Teaching Models to Improve on TapeLiat Bezalel, Eyal Orgad, Amir GlobersonAAAI 2025
- MallowsPO: Fine-Tune Your LLM with Preference DispersionsHaoxian Chen, Hanyang Zhao, Henry Lam, David D. Yao et al.ICLR 2025 · 1 citation
- Improving Context-Aware Preference Modeling for Language ModelsSilviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro SordoniNeurIPS 2024 · 30 citations
