User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
Yuhan Liu, Michael J. Q. Zhang, Eunsol Choi
Abstract
Once language models (LMs) are deployed, they can interact with users long-term, ideally evolving based on their feedback. Asking for direct user feedback can be disruptive; thus, we study harvesting implicit user feedback from user-LM interaction logs. We study two user-LM interaction datasets (WildChat and LMSYS). First, we analyze user feedback in the user-LLM conversation logs, providing insights into when and why such feedback occurs. Second, we study harvesting learning signals from such implicit user feedback. Specifically, we study whether incorporating the contents of user feedback (e.g., user wanted clarification), in addition to the polarity of the feedback, can improve the model performance. We observe mixed results, showing this helps in short human-designed questions (MTBench) but not on longer and more complex questions (Wild-Bench). Together, we provide an in-depth study of implicit user feedback, showing its potential and limitations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73b7b50f-c6eb-4ede-80e4-e5a5a1a27aa2Cited by top-tier papers6
- PrefDisco: Benchmarking Proactive Personalized ReasoningShuyue Stella Li, Avinandan Bose, Faeze Brahman, Simon S. Du et al.ICLR 2026 · 8 citations
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference LearningYifan Wang, Bolian Li, Junlin Wu, Zhaoxuan Tan et al.ICLR 2026 · 5 citations
- WildReward: Learning Reward Models from In-the-Wild Human InteractionsHao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao et al.ACL 2026 · 3 citations
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen et al.ICML 2026 · 2 citations
- ReIn: Conversational Error Recovery with Reasoning InceptionTakyoung Kim, Jinseok Nam, Chandrayee Basu, Xing Fan et al.ICLR 2026 · 1 citation
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 491 citations
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
Related papers
- WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin et al.ACL 2026 · 35 citations
- Let Me Ask You This: How Can a Voice Assistant Elicit Explicit User Feedback?Ziang Xiao, Sarah Mennicken, Bernd Huber, Adam Shonkoff et al.CSCW 2021 · 11 citations
- Interpretable User Satisfaction Estimation for Conversational Systems with Large Language ModelsYing-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang et al.ACL 2024 · 11 citations
- A Scalable Framework for Learning From Implicit User Feedback to Improve Natural Language Understanding in Large-Scale Conversational AI SystemsSunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal et al.EMNLP 2021 · 17 citations
- Retrospective Learning from InteractionsZizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng et al.ACL 2025
