Lune

EMNLP2024顶会

Towards Aligning Language Models with Textual Feedback

Saüc Abadal Lloret, Shehzaad Dhuliawala, Keerthiram Murugesan, Mrinmaya Sachan

2024年份
2被引次数
3顶会引用

摘要

We present ALT (ALignment with Textual feedback), an approach that aligns language models with user preferences expressed in text. We argue that text offers greater expressiveness, enabling users to provide richer feedback than simple comparative preferences, leading to more efficient and effective alignment. ALT aligns the model by conditioning its generations on the textual feedback. Our method relies solely on language modeling techniques and requires minimal hyper-parameter tuning while retaining the main benefits of RL-based alignment algorithms. We demonstrate the efficacy and efficiency of textual feedback across different tasks, including toxicity reduction, summarization, and dialogue response generation. Notably, ALT outperforms PPO in toxicity reduction and matches its performance on summarization with only 20% of the samples. We also explore using ALT with feedback from an existing LLM, examining constrained and unconstrained feedback. Additionally, we outline future directions to align models with natural language feedback. 1

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖