Lune

EMNLP2024Top-tier venue

Towards Aligning Language Models with Textual Feedback

Saüc Abadal Lloret, Shehzaad Dhuliawala, Keerthiram Murugesan, Mrinmaya Sachan

2024Year
2Citations
3Top-tier citations

Abstract

We present ALT (ALignment with Textual feedback), an approach that aligns language models with user preferences expressed in text. We argue that text offers greater expressiveness, enabling users to provide richer feedback than simple comparative preferences, leading to more efficient and effective alignment. ALT aligns the model by conditioning its generations on the textual feedback. Our method relies solely on language modeling techniques and requires minimal hyper-parameter tuning while retaining the main benefits of RL-based alignment algorithms. We demonstrate the efficacy and efficiency of textual feedback across different tasks, including toxicity reduction, summarization, and dialogue response generation. Notably, ALT outperforms PPO in toxicity reduction and matches its performance on summarization with only 20% of the samples. We also explore using ALT with feedback from an existing LLM, examining constrained and unconstrained feedback. Additionally, we outline future directions to align models with natural language feedback. 1

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers3

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines