RLVF: Learning from Verbal Feedback without Overgeneralization
Moritz Stephan, Alexander Khazatsky, Eric Mitchell, Annie S. Chen, Sheryl Hsu, Archit Sharma, Chelsea Finn
摘要
The diversity of contexts in which large language models (LLMs) are deployed requires the ability to modify or customize default model behaviors to incorporate nuanced requirements and preferences. A convenient interface to specify such model adjustments is high-level verbal feedback, such as"Don't use emojis when drafting emails to my boss."However, while writing high-level feedback is far simpler than collecting annotations for reinforcement learning from human feedback (RLHF), we find that simply prompting a model with such feedback leads to overgeneralization of the feedback to contexts where it is not relevant. We study the problem of incorporating verbal feedback without such overgeneralization, inspiring a new method Contextualized Critiques with Constrained Preference Optimization (C3PO). C3PO uses a piece of high-level feedback to generate a small synthetic preference dataset specifying how the feedback should (and should not) be applied. It then fine-tunes the model in accordance with the synthetic preference data while minimizing the divergence from the original model for prompts where the feedback does not apply. Our experimental results indicate that our approach effectively applies verbal feedback to relevant scenarios while preserving existing behaviors for other contexts. For both human- and GPT-4-generated high-level feedback, C3PO effectively adheres to the given feedback comparably to in-context baselines while reducing overgeneralization by 30%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Characterizing Photorealism and Artifacts in Diffusion Model-Generated ImagesNegar Kamali, Karyn Nakamura, Aakriti Kumar, Angelos Chatzimparmpas 等CHI 2025 · 被引用 23 次
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann 等ICML 2026
- Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasNishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng 等ACL 2025
- Let Me Teach You: Pedagogical Foundations of Feedback for Language ModelsBeatriz Borges, Niket Tandon, Tanja Käser, Antoine BosselutEMNLP 2024
它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 等ICML 2022 · 被引用 520 次
相关 Paper
- Rule Based Rewards for Language Model SafetyTong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam 等NeurIPS 2024 · 被引用 159 次
- Explicit Preference Optimization: No Need for an Implicit Reward ModelXiangkun Hu, Lemin Kong, Tong He, David WipfICML 2025
- Teaching Models to Improve on TapeLiat Bezalel, Eyal Orgad, Amir GlobersonAAAI 2025
- MallowsPO: Fine-Tune Your LLM with Preference DispersionsHaoxian Chen, Hanyang Zhao, Henry Lam, David D. Yao 等ICLR 2025 · 被引用 1 次
- Improving Context-Aware Preference Modeling for Language ModelsSilviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro SordoniNeurIPS 2024 · 被引用 30 次
