Aligning LLM Agents by Learning Latent Preference from User Edits
Ge Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro, Dipendra Misra
Abstract
We study interactive learning of LLM-based language agents based on user edits made to the agent's output. In a typical setting such as writing assistants, the user interacts with a language agent to generate a response given a context, and may optionally edit the agent response to personalize it based on their latent preference, in addition to improving the correctness. The edit feedback is naturally generated, making it a suitable candidate for improving the agent's alignment with the user's preference, and for reducing the cost of user edits over time. We propose a learning framework, PRELUDE that infers a description of the user's latent preference based on historic edit data. The inferred user preference descriptions are used to define prompts for generating responses in the future. This avoids fine-tuning the agent, which is costly, challenging to scale with the number of users, and may even degrade its performance on other tasks. Furthermore, learning descriptive preference improves interpretability, allowing the user to view and modify the learned preference. However, user preference can be complex, subtle, and vary based on context, making it challenging to learn. To address this, we propose a simple yet effective algorithm named CIPHER that leverages the LLM to infer the user preference for a given context based on user edits. In the future, CIPHER retrieves inferred preferences from the k-closest contexts in the history, and forms an aggregate preference for response generation. We introduce two interactive environments -- summarization and email writing, and use a GPT-4 simulated user for evaluation. On both tasks, CIPHER outperforms several baselines by achieving the lowest edit distance cost while only having a small overhead in LLM query cost. Our analysis reports that user preferences learned by CIPHER show significant similarity to the ground truth latent preferences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live InteractionsBufang Yang, Yunqi Guo, Lilin Xu, Zhenyu Yan et al.UbiComp 2025 · 26 citations
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 14 citations
- User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning SignalYuhan Liu, Michael J. Q. Zhang, Eunsol ChoiEMNLP 2025 · 12 citations
- Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric AssessmentsMarharyta Domnich, Julius Välja, Rasmus Moorits Veski, Giacomo Magnifico et al.AAAI 2025 · 11 citations
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference LearningYifan Wang, Bolian Li, Junlin Wu, Zhaoxuan Tan et al.ICLR 2026 · 5 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li et al.ASE 2022 · 81 citations
Related papers
- Aligning LLMs by Predicting Preferences from User Writing SamplesStéphane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald, Katherine MetcalfICML 2025
- Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and RewardDipendra Misra, Aldo Pacchiano, Ta-Chung Chi, Ge GaoNeurIPS 2025
- Learning to summarize user information for personalized reinforcement learning from human feedbackHyunJi Nam, Yanming Wan, Mickel Liu, Peter F. Ahnn et al.ICLR 2026 · 10 citations
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMsSiyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika et al.ICLR 2025
- Adaptive Preference Arithmetic: A Personalized Agent with Adaptive Preference Arithmetic for Dynamic Preference ModelingHongyi Nie, Yaqing Wang, Mingyang Zhou, Feiyang Pan et al.NeurIPS 2025 · 1 citation
