Towards Aligning Language Models with Textual Feedback
Saüc Abadal Lloret, Shehzaad Dhuliawala, Keerthiram Murugesan, Mrinmaya Sachan
摘要
We present ALT (ALignment with Textual feedback), an approach that aligns language models with user preferences expressed in text. We argue that text offers greater expressiveness, enabling users to provide richer feedback than simple comparative preferences, leading to more efficient and effective alignment. ALT aligns the model by conditioning its generations on the textual feedback. Our method relies solely on language modeling techniques and requires minimal hyper-parameter tuning while retaining the main benefits of RL-based alignment algorithms. We demonstrate the efficacy and efficiency of textual feedback across different tasks, including toxicity reduction, summarization, and dialogue response generation. Notably, ALT outperforms PPO in toxicity reduction and matches its performance on summarization with only 20% of the samples. We also explore using ALT with feedback from an existing LLM, examining constrained and unconstrained feedback. Additionally, we outline future directions to align models with natural language feedback. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Recontextualization Mitigates Specification Gaming Without Modifying the SpecificationAriana Azarbal, Victor Gillioz, Vladimir Ivanov, Bryce Woodworth 等ICML 2026 · 被引用 11 次
- RL with Learnable Textual Feedback: A Bilevel ApproachUtsav Singh, Sidhaarth Murali, Souradip Chakraborty, Amrit Singh BediICML 2026 · 被引用 1 次
- Reward-Augmented Data Enhances Direct Preference Alignment of LLMsShenao Zhang, Zhihan Liu, Boyi Liu, Yufeng Zhang 等ICML 2025
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and FeedbackSongyang Gao, Qiming Ge, Wei Shen, Shihan Dou 等ICML 2024 · 被引用 24 次
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao 等ICML 2023 · 被引用 287 次
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar 等ICML 2024 · 被引用 212 次
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang 等ICML 2024 · 被引用 136 次
- Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual FeedbackYafu Li, Xuyang Hu, Xiaoye Qu, Linjie Li 等ICML 2025
