ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
Xinyi Hu, Yuhao Shen, Zhang Baolin, Hengxin Zhang, Jun Dai, Shuang Ge, Yue Li, Mingcheng Wan
摘要
Speculative Decodin promises to accelerate Large Language Model inference, yet its efficacy often degrades in production-grade scenarios. Existing evaluations typically overlook the compute-bound nature of high-concurrency regimes, where verification compute becomes the dominant bottleneck. Consequently, prior methods face a dilemma: static trees incur massive verification waste, while dynamic trees suffer from cumulative misjudgments and kernel incompatibility. To bridge this gap, we introduce ECHO, a high concurrency-oriented framework integrated into SGLang that reformulates speculative execution as a budgeted scheduling problem. Crucially, ECHO employs sparse confidence gating to manage the batch as a unified super-tree, elastically pivoting budget between depth and width to co-optimize the trade-off between reducing global verification steps and maximizing per-step efficiency. Extensive evaluations across diverse model scales—particularly the industrial-grade Qwen3-235B—demonstrate that ECHO consistently outperforms state-of-the-art baselines in both low-load and high-load scenarios, achieving up to 5.35 walltime speedup and delivering over 20% relative speedup gain against the strongest baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 被引用 347 次
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 143 次
相关 Paper
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative DecodingShenggui Li, Chao Wang, YIKAI ZHU, Yubo Wang 等ICML 2026 · 被引用 7 次
- Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative ExplorationShuzhang Zhong, Haochen Huang, Shengxuan Qiu, Pengfei Zuo 等OSDI 2026
- HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative DecodingSiran Liu, Yang Ye, Qianchao Zhu, Zane Cao 等ACL 2026 · 被引用 2 次
- ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated VerificationSiran Liu, Zane Cao, Yongchao HeACL 2026 · 被引用 3 次
- Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative DecodingWen-Hung Lee, Jian-Jia Chen, Xiaolin Lin, Pei-Shuo Wang 等ICML 2026
