Speculative Actions: A Lossless Framework for Faster AI Agents
Naimeng Ye, Arnav Ahuja, Georgios Liargkovas, Yunan Lu, Kostis Kaffes, Tianyi Peng
摘要
AI agents are increasingly deployed in complex, interactive environments, yet their runtime remains a major bottleneck for training, evaluation, and real-world use. Typical agent behavior unfolds sequentially, with each action requiring an API call that can incur substantial latency. For example, a game of chess between two state-of-the-art agents can take hours. We introduce speculative actions, a lossless acceleration framework for general agentic systems. Inspired by speculative execution in microprocessors and speculative decoding in LLM inference, our method uses faster models to predict likely future actions and execute them in parallel, committing only when predictions match. We evaluate speculative actions across gaming, e-commerce, and web search environments, and additionally study a lossy extension in an operating systems setting. Across domains, we achieve up to 55% next-action prediction accuracy, translating into up to 20% latency reductions. Finally, we present a cost-latency analysis that formalizes the tradeoff between speculative breadth and time savings. This analysis enables principled tuning and selective branch launching, to ensure that multi-branch speculation delivers practical speedups without prohibitive cost growth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionDongfu Jiang, Xiang Ren, Bill Yuchen LinACL 2023 · 被引用 95 次
- Scaling Speculative Decoding with Lookahead ReasoningYichao Fu, Rui Ge, Zelei Shao, Zhijie Deng 等NeurIPS 2025 · 被引用 14 次
- ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT PipelinesTengjun Jin, Yuxuan Zhu, Daniel KangVLDB 2026 · 被引用 13 次
- Dynamic Speculative Agent PlanningYilin Guan, Qingfeng Lan, Fei Sun, Dujian Ding 等ICLR 2026 · 被引用 10 次
相关 Paper
- SuffixDecoding: Extreme Speculative Decoding for Emerging AI ApplicationsGabriele Oliaro, Zhihao Jia, Daniel F. Campos, Aurick QiaoNeurIPS 2025 · 被引用 34 次
- Speculative Monte-Carlo Tree SearchScott Cheng, Mahmut T. Kandemir, Ding-Yong HongNeurIPS 2024 · 被引用 4 次
- Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed AcceptanceSongsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu 等EMNLP 2025
- Re-SpS: A Reinforcement Learning Approach to Speculative SamplingChenan Wang, Daniel H. Shi, Haipeng ChenAAAI 2026
- EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU UtilizationYize Wu, Ke Gao, Ling Li, Yanjun WuNeurIPS 2025 · 被引用 3 次
