Substance over Style: Evaluating Proactive Conversational Coaching Agents
Vidya Srinivas, Xuhai Xu, Xin Liu, Kumar Ayush, Isaac R. Galatzer-Levy, Shwetak N. Patel, Daniel McDuff, Tim Althoff
Abstract
While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initially undefined goals that evolve through multi-turn interactions, subjective evaluation criteria, and mixed-initiative dialogue. In this work, we describe and implement five multi-turn coaching agents that exhibit distinct conversational styles, and evaluate them through a user study, collecting first-person feedback on 155 conversations. We find that users highly value core functionality, and that stylistic components in absence of core components are viewed negatively. By comparing user feedback with thirdperson evaluations from health experts and an LM, we reveal significant misalignment across evaluation approaches. Our findings provide insights into design and evaluation of conversational coaching agents and contribute toward improving human-centered NLP applications. * Work done during an internship at Google There's quite a few factors. I'll get home after work, and after dinner and other responsibilities, I'm super tired. I'm also a DJ and I have work as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45da9e32-20b7-4d9d-a10e-38bdf60e9434Cited by top-tier papers3
- Bloom: Designing for LLM-Augmented Behavior Change InteractionsMatthew Jörke, Defne Genç, Valentin Teutschbein, Shardul Sapkota et al.CHI 2026 · 6 citations
- Toward Flexible Psychiatric History-Taking and Visualization: Exploring Clinician Perspectives with Large Language ModelsYugyeong Jung, Thu Hoang Anh Vo, Hyun Seung Moon, Jae Young Choi et al.CHI 2026 · 1 citation
- Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-gated LLMLuo Ji, Qi Qin, Ningyuan Xi, Teng Chen et al.ICML 2026
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Facilitating Self-Guided Mental Health Interventions Through Human-Language Model Interaction: A Case Study of Cognitive RestructuringAshish Sharma, Kevin Rushton, Inna Wanyin Lin, Theresa Nguyen et al.CHI 2024 · 104 citations
- Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMsZhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao et al.NeurIPS 2024 · 43 citations
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He et al.ACL 2024 · 35 citations
- Cognitive Reframing of Negative Thoughts through Human-Language Model InteractionAshish Sharma, Kevin Rushton, Inna E. Lin, David Wadden et al.ACL 2023 · 32 citations
Related papers
- Efficient Management of LLM-Based Coaching Agents' Reasoning While Maintaining Interaction Quality and SpeedAndreas Göldi, Roman Rietsche, Lyle H. UngarCHI 2025 · 6 citations
- "Having Lunch Now": Understanding How Users Engage with a Proactive Agent for Daily Planning and Self-ReflectionAdnan Abbas, Caleb Wohn, Arnav Jagtap, Eugenia Ha Rim Rho et al.CHI 2026 · 1 citation
- T2 Coach: A Qualitative Study of an Automated Health Coach for Diabetes Self-ManagementElliot G. Mitchell, Pooja M. Desai, Arlene M. Smaldone, Andrea Cassells et al.CHI 2025 · 10 citations
- GPTCoach: Towards LLM-Based Physical Activity CoachingMatthew Jörke, Shardul Sapkota, Lyndsea Warkenthien, Niklas Vainio et al.CHI 2025 · 89 citations
- Understanding How eHealth Coaches Tailor Support For Weight Loss: Towards the Design of Person-Centered Coaching SystemsKathleen Ryan, Samantha Dockray, Conor LinehanCHI 2022 · 26 citations
