Tvcache: A Tool-Value Cache for Post-Training LLM Agents
Abhishek Vijaya Kumar, Bhaskar Kataria, Byungsoo Oh, Emaad Manzoor, Rachee Singh
摘要
In RL post-training of LLM agents, calls to external tools take several seconds or even minutes, leaving allocated GPUs idle and inflating post-training time and cost. While many tool invocations repeat across parallel rollouts and could in principle be cached, naively caching their outputs for reuse is incorrect since tool outputs depend on the environment state induced by prior agent interactions. We present TVCACHE, a stateful tool-value cache for LLM agent post-training. TVCACHE maintains a tree of observed tool-call sequences and performs longest-prefix matching for cache lookups: a hit occurs only when the agent’s full tool history matches a previously executed sequence, guaranteeing identical environment state. On three diverse workloads—terminal-based tasks, SQL generation, and video understanding—TVCACHE achieves cache hit rates of up to 70% and reduces median tool call execution time by up to 6.9×, with no degradation in post-training reward accumulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang 等ICLR 2026 · 被引用 406 次
- KVLink: Accelerating Large Language Models via Efficient KV Cache ReuseJingbo Yang, Bairu Hou, Wei Wei, Yujia Bao 等NeurIPS 2025 · 被引用 83 次
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced ReasoningShaokun Zhang, Yi Dong, Jieyu Zhang, Jan Kautz 等ICLR 2026 · 被引用 61 次
相关 Paper
- Parallelizing LLM Agent Execution with Contrastive Task AllocationYuyang Peng, Yanling Xu, Shuyi Wang, Xiaofei Liao 等KDD 2026 · 被引用 1 次
- AgentVocab: Structure-Aware Vocabulary Adaptation for Efficient LLM AgentsKai Bian, Haosi Mo, Xuebo Liu, Shuangyong Song 等ICML 2026
- Generative Caching for Structurally Similar Prompts and ResponsesSarthak Chakraborty, Suman Nath, Xuchao Zhang, Chetan Bansal 等NeurIPS 2025 · 被引用 5 次
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu 等NeurIPS 2025 · 被引用 77 次
- AReaL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language ModelsJiarui Zhang, Yuchen Yang, Ran Yan, Zhiyu Mei 等ICML 2026
