IEvoAgent: Evolving Conversational Agent based on User Implicit Feedback
Yichen Cai, Jiayang Li, Junyuan Qiu, Jingya Guo, Weitao You, Changyuan Yang, Lingyun Sun, Pei Chen
摘要
Current conversational agents often follow static learning paradigms and miss the implicit, evolving feedback embedded in users' follow-up behaviors. We propose IEvoAgent, an evolving conversational agent framework that leverages the structured dependency between agent responses and user reactions. We construct an annotated dataset from LMSYS-Chat-1M and WildChat and find consistent response-conditioned feedback patterns. Based on this finding, IEvoAgent uses a conditional feedback distribution matrix to estimate expected feedback rewards, combining offline KTO alignment with an inference-time promptevolution mechanism driven by a dynamic matrix. Experiments on MT-Bench-101, Wild-Bench, and FB-Bench show improvements over open-source baselines, indicating that mining implicit feedback supports better multiturn alignment under evolving user preferences. Our code and dataset are available at https://github.com/Hualeez/IEvoAgent .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He 等ACL 2024 · 被引用 35 次
相关 Paper
- WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin 等ACL 2026 · 被引用 35 次
- User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning SignalYuhan Liu, Michael J. Q. Zhang, Eunsol ChoiEMNLP 2025 · 被引用 12 次
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMsSiyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika 等ICLR 2025
- Aligning Language Models Using Follow-up Likelihood as Reward SignalChen Zhang, Dading Chong, Feng Jiang, Chengguang Tang 等AAAI 2025 · 被引用 6 次
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question AnsweringSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasACL 2026 · 被引用 5 次
