Efficient Management of LLM-Based Coaching Agents' Reasoning While Maintaining Interaction Quality and Speed
Andreas Göldi, Roman Rietsche, Lyle H. Ungar
摘要
LLM-based agents improve upon standalone LLMs, which are optimized for immediate intent-satisfaction, by allowing the pursuit of more extended objectives, such as helping users over the long term. To do so, LLM-based agents need to reason before responding. For complex tasks like personalized coaching, this reasoning can be informed by adding relevant information at key moments, shifting it in the desired direction. However, the pursuit of objectives beyond interaction quality may compromise this very quality. Moreover, as the depth and informativeness of reasoning increase, so do the number of tokens required, leading to higher latency and cost. This study investigates how an LLM-based coaching agent can adjust its reasoning depth using a discrepancy mechanism that signals how much reasoning effort to allocate based on how well the objective is being met. Our discrepancy-based mechanism constrains reasoning to better align with alternative objectives, reducing cost roughly tenfold while minimally impacting interaction quality.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Learning to Reason Efficiently with Discounted Reinforcement LearningAlex Ayoub, Kavosh Asadi, Dale Schuurmans, Csaba Szepesvari 等ICLR 2026 · 被引用 4 次
- DEPO: Dual-Efficiency Preference Optimization for LLM AgentsSirui Chen, Mengshi Zhao, Lei Xu, Yuying Zhao 等AAAI 2026 · 被引用 2 次
- Routing, Cascades, and User Choice for LLMsRafid MahmoodICLR 2026 · 被引用 2 次
- Reasoning Can Be Restored by Correcting a Few Decision TokensShen Changshuo, Leheng Sheng, Yuxin Chen, Xiang Wang 等ICML 2026 · 被引用 2 次
- CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal ControlZirui Yuan, Siqi Lai, Hao LiuICLR 2026 · 被引用 18 次
