Efficient Management of LLM-Based Coaching Agents' Reasoning While Maintaining Interaction Quality and Speed
Andreas Göldi, Roman Rietsche, Lyle H. Ungar
Abstract
LLM-based agents improve upon standalone LLMs, which are optimized for immediate intent-satisfaction, by allowing the pursuit of more extended objectives, such as helping users over the long term. To do so, LLM-based agents need to reason before responding. For complex tasks like personalized coaching, this reasoning can be informed by adding relevant information at key moments, shifting it in the desired direction. However, the pursuit of objectives beyond interaction quality may compromise this very quality. Moreover, as the depth and informativeness of reasoning increase, so do the number of tokens required, leading to higher latency and cost. This study investigates how an LLM-based coaching agent can adjust its reasoning depth using a discrepancy mechanism that signals how much reasoning effort to allocate based on how well the objective is being met. Our discrepancy-based mechanism constrains reasoning to better align with alternative objectives, reducing cost roughly tenfold while minimally impacting interaction quality.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Learning to Reason Efficiently with Discounted Reinforcement LearningAlex Ayoub, Kavosh Asadi, Dale Schuurmans, Csaba Szepesvari et al.ICLR 2026 · 4 citations
- DEPO: Dual-Efficiency Preference Optimization for LLM AgentsSirui Chen, Mengshi Zhao, Lei Xu, Yuying Zhao et al.AAAI 2026 · 2 citations
- Routing, Cascades, and User Choice for LLMsRafid MahmoodICLR 2026 · 2 citations
- Reasoning Can Be Restored by Correcting a Few Decision TokensShen Changshuo, Leheng Sheng, Yuxin Chen, Xiang Wang et al.ICML 2026 · 2 citations
- CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal ControlZirui Yuan, Siqi Lai, Hao LiuICLR 2026 · 18 citations
