Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models
Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar
Abstract
Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational patterns in both general-purpose (Chat-GPT and Bing Copilot) and task-oriented (customer service chatbot) conversational systems. Existing approaches based on featurized ML models or text embeddings fall short in extracting generalizable patterns and are hard to interpret. In this work, we show that LLMs can extract interpretable signals of user satisfaction from their natural language utterances more effectively than embedding-based approaches. Moreover, an LLM can be tailored for USE via an iterative prompting framework using supervision from labeled examples. Our proposed method, Supervised Prompting for User satisfaction Rubrics (SPUR), not only has higher accuracy but is more interpretable as it scores user satisfaction via learned rubrics with a detailed breakdown.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e960556-174b-45cf-9f76-493926ab10f5Cited by top-tier papers6
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- Conversation Progress Guide : UI System for Enhancing Self-Efficacy in Conversational AIDaeun Jeong, Sungbok Shin, Jongwook JeongCHI 2025 · 12 citations
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference LearningYifan Wang, Bolian Li, Junlin Wu, Zhaoxuan Tan et al.ICLR 2026 · 5 citations
- Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the WildSheshera Mysore, Debarati Das, Hancheng Cao, Bahareh SarrafzadehEMNLP 2025 · 3 citations
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu et al.ICML 2026
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
- User Satisfaction Estimation with Sequential Dialogue Act Modeling in Goal-oriented Conversational SystemsYang Deng, Wenxuan Zhang, Wai Lam, Hong Cheng et al.WWW 2022 · 34 citations
- Modeling User Satisfaction Dynamics in Dialogue via Hawkes ProcessFanghua Ye, Zhiyuan Hu, Emine YilmazACL 2023 · 6 citations
Related papers
- Transparent and Scrutable Recommendations Using Natural Language User ProfilesJerome Ramos, Hossein A. Rahmani, Xi Wang, Xiao Fu et al.ACL 2024
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- A Scalable Framework for Learning From Implicit User Feedback to Improve Natural Language Understanding in Large-Scale Conversational AI SystemsSunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal et al.EMNLP 2021 · 17 citations
- Measuring Intent Comprehension in LLMsNadav Kunievsky, James EvansICML 2026 · 1 citation
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language ModelsXiaolei Wang, Xinyu Tang, Xin Zhao, Jingyuan Wang et al.EMNLP 2023 · 69 citations
