Substance over Style: Evaluating Proactive Conversational Coaching Agents
Vidya Srinivas, Xuhai Xu, Xin Liu, Kumar Ayush, Isaac R. Galatzer-Levy, Shwetak N. Patel, Daniel McDuff, Tim Althoff
摘要
While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initially undefined goals that evolve through multi-turn interactions, subjective evaluation criteria, and mixed-initiative dialogue. In this work, we describe and implement five multi-turn coaching agents that exhibit distinct conversational styles, and evaluate them through a user study, collecting first-person feedback on 155 conversations. We find that users highly value core functionality, and that stylistic components in absence of core components are viewed negatively. By comparing user feedback with thirdperson evaluations from health experts and an LM, we reveal significant misalignment across evaluation approaches. Our findings provide insights into design and evaluation of conversational coaching agents and contribute toward improving human-centered NLP applications. * Work done during an internship at Google There's quite a few factors. I'll get home after work, and after dinner and other responsibilities, I'm super tired. I'm also a DJ and I have work as well.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bloom: Designing for LLM-Augmented Behavior Change InteractionsMatthew Jörke, Defne Genç, Valentin Teutschbein, Shardul Sapkota 等CHI 2026 · 被引用 6 次
- Toward Flexible Psychiatric History-Taking and Visualization: Exploring Clinician Perspectives with Large Language ModelsYugyeong Jung, Thu Hoang Anh Vo, Hyun Seung Moon, Jae Young Choi 等CHI 2026 · 被引用 1 次
- Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-gated LLMLuo Ji, Qi Qin, Ningyuan Xi, Teng Chen 等ICML 2026
它引用的顶会 Paper9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Facilitating Self-Guided Mental Health Interventions Through Human-Language Model Interaction: A Case Study of Cognitive RestructuringAshish Sharma, Kevin Rushton, Inna Wanyin Lin, Theresa Nguyen 等CHI 2024 · 被引用 104 次
- Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMsZhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao 等NeurIPS 2024 · 被引用 43 次
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He 等ACL 2024 · 被引用 35 次
- Cognitive Reframing of Negative Thoughts through Human-Language Model InteractionAshish Sharma, Kevin Rushton, Inna E. Lin, David Wadden 等ACL 2023 · 被引用 32 次
相关 Paper
- Efficient Management of LLM-Based Coaching Agents' Reasoning While Maintaining Interaction Quality and SpeedAndreas Göldi, Roman Rietsche, Lyle H. UngarCHI 2025 · 被引用 6 次
- "Having Lunch Now": Understanding How Users Engage with a Proactive Agent for Daily Planning and Self-ReflectionAdnan Abbas, Caleb Wohn, Arnav Jagtap, Eugenia Ha Rim Rho 等CHI 2026 · 被引用 1 次
- T2 Coach: A Qualitative Study of an Automated Health Coach for Diabetes Self-ManagementElliot G. Mitchell, Pooja M. Desai, Arlene M. Smaldone, Andrea Cassells 等CHI 2025 · 被引用 10 次
- GPTCoach: Towards LLM-Based Physical Activity CoachingMatthew Jörke, Shardul Sapkota, Lyndsea Warkenthien, Niklas Vainio 等CHI 2025 · 被引用 89 次
- Understanding How eHealth Coaches Tailor Support For Weight Loss: Towards the Design of Person-Centered Coaching SystemsKathleen Ryan, Samantha Dockray, Conor LinehanCHI 2022 · 被引用 26 次
