Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models
Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar
摘要
Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational patterns in both general-purpose (Chat-GPT and Bing Copilot) and task-oriented (customer service chatbot) conversational systems. Existing approaches based on featurized ML models or text embeddings fall short in extracting generalizable patterns and are hard to interpret. In this work, we show that LLMs can extract interpretable signals of user satisfaction from their natural language utterances more effectively than embedding-based approaches. Moreover, an LLM can be tailored for USE via an iterative prompting framework using supervision from labeled examples. Our proposed method, Supervised Prompting for User satisfaction Rubrics (SPUR), not only has higher accuracy but is more interpretable as it scores user satisfaction via learned rubrics with a detailed breakdown.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme 等ACL 2024 · 被引用 27 次
- Conversation Progress Guide : UI System for Enhancing Self-Efficacy in Conversational AIDaeun Jeong, Sungbok Shin, Jongwook JeongCHI 2025 · 被引用 12 次
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference LearningYifan Wang, Bolian Li, Junlin Wu, Zhaoxuan Tan 等ICLR 2026 · 被引用 5 次
- Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the WildSheshera Mysore, Debarati Das, Hancheng Cao, Bahareh SarrafzadehEMNLP 2025 · 被引用 3 次
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu 等ICML 2026
它引用的顶会 Paper7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- User Satisfaction Estimation with Sequential Dialogue Act Modeling in Goal-oriented Conversational SystemsYang Deng, Wenxuan Zhang, Wai Lam, Hong Cheng 等WWW 2022 · 被引用 34 次
- Modeling User Satisfaction Dynamics in Dialogue via Hawkes ProcessFanghua Ye, Zhiyuan Hu, Emine YilmazACL 2023 · 被引用 6 次
相关 Paper
- Transparent and Scrutable Recommendations Using Natural Language User ProfilesJerome Ramos, Hossein A. Rahmani, Xi Wang, Xiao Fu 等ACL 2024
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
- A Scalable Framework for Learning From Implicit User Feedback to Improve Natural Language Understanding in Large-Scale Conversational AI SystemsSunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal 等EMNLP 2021 · 被引用 17 次
- Measuring Intent Comprehension in LLMsNadav Kunievsky, James EvansICML 2026 · 被引用 1 次
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language ModelsXiaolei Wang, Xinyu Tang, Xin Zhao, Jingyuan Wang 等EMNLP 2023 · 被引用 69 次
