User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
Yuhan Liu, Michael J. Q. Zhang, Eunsol Choi
摘要
Once language models (LMs) are deployed, they can interact with users long-term, ideally evolving based on their feedback. Asking for direct user feedback can be disruptive; thus, we study harvesting implicit user feedback from user-LM interaction logs. We study two user-LM interaction datasets (WildChat and LMSYS). First, we analyze user feedback in the user-LLM conversation logs, providing insights into when and why such feedback occurs. Second, we study harvesting learning signals from such implicit user feedback. Specifically, we study whether incorporating the contents of user feedback (e.g., user wanted clarification), in addition to the polarity of the feedback, can improve the model performance. We observe mixed results, showing this helps in short human-designed questions (MTBench) but not on longer and more complex questions (Wild-Bench). Together, we provide an in-depth study of implicit user feedback, showing its potential and limitations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- PrefDisco: Benchmarking Proactive Personalized ReasoningShuyue Stella Li, Avinandan Bose, Faeze Brahman, Simon S. Du 等ICLR 2026 · 被引用 8 次
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference LearningYifan Wang, Bolian Li, Junlin Wu, Zhaoxuan Tan 等ICLR 2026 · 被引用 5 次
- WildReward: Learning Reward Models from In-the-Wild Human InteractionsHao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao 等ACL 2026 · 被引用 3 次
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen 等ICML 2026 · 被引用 2 次
- ReIn: Conversational Error Recovery with Reasoning InceptionTakyoung Kim, Jinseok Nam, Chandrayee Basu, Xing Fan 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 被引用 491 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen 等ICLR 2024 · 被引用 308 次
相关 Paper
- WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin 等ACL 2026 · 被引用 35 次
- Let Me Ask You This: How Can a Voice Assistant Elicit Explicit User Feedback?Ziang Xiao, Sarah Mennicken, Bernd Huber, Adam Shonkoff 等CSCW 2021 · 被引用 11 次
- Interpretable User Satisfaction Estimation for Conversational Systems with Large Language ModelsYing-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang 等ACL 2024 · 被引用 11 次
- A Scalable Framework for Learning From Implicit User Feedback to Improve Natural Language Understanding in Large-Scale Conversational AI SystemsSunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal 等EMNLP 2021 · 被引用 17 次
- Retrospective Learning from InteractionsZizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng 等ACL 2025
