E^2-SCI: Elastic Edge–Cloud Speculative Decoding via Credit Inertia
Senyao Li, Haozhao Wang, Zhaobai Jiang, Zhanbo Jin, Hao Fan, Ruixuan Li
Abstract
In edge–cloud environments, the efficiency of speculative decoding is heavily constrained by uplink transmission and cloud-side verification. In this work, we identify a phenomenon we term credit inertia, where the acceptance rates of adjacent token windows exhibit strong temporal consistency. Tokens following recently well-performing windows are likely to pass verification, whereas tokens following poorly performing windows are likely to fail. Motivated by this observation, we propose E-SCI, an elastic edge–cloud speculative decoding framework that dynamically adjusts draft token verification thresholds based on recent historical performance. This adaptive mechanism allows the system to be more permissive for windows with strong historical performance and stricter for windows with weak performance, effectively leveraging temporal consistency to reduce overall latency. We further introduce Progressive Lookahead Concurrency (PLC), which pipelines draft generation and verification asynchronously to hide latency. Experiments across multiple benchmarks show that E-SCI achieves over tokens/s on DeepSeek-R1-Distill-Qwen (1.5B/32B), delivering an 88.5% speed improvement over the FSD baseline while maintaining accuracy. Notably, E-SCI integrates seamlessly with existing frameworks (e.g., EAGLE-3), demonstrating broad applicability and superior efficiency–quality trade-offs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fd57421-0932-46e1-b072-1f75ee15b449Builds on23
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 347 citations
- GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative DecodingCunxiao Du, Jing Jiang, Yuanchen Xu, Jiawei Wu et al.ICML 2024 · 72 citations
- SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer DevicesRuslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen et al.NeurIPS 2024 · 70 citations
- Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe SamplingYiran Zhao, Wenyue Zheng, Tianle Cai, Do Xuan Long et al.NeurIPS 2024 · 46 citations
Related papers
- PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative DecodingHan Yunhe, Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi et al.ICML 2026 · 2 citations
- Overcoming Joint Intractability with Lossless Hierarchical Speculative DecodingYuxuan Zhou, Fei Huang, Heng Li, Fengyi Wu et al.ICLR 2026 · 5 citations
- HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative DecodingSiran Liu, Yang Ye, Qianchao Zhu, Zane Cao et al.ACL 2026 · 2 citations
- HCSpec: Two-Tier Horizontal Cascade Speculative Decoding for High-Efficiency Large Language Model InferenceYizhou Zhang, Siming Chen, Hao Ye, Erhu FengACL 2026
- EdgeSpec: Distributed Speculative Decoding for Large Language Models at EdgeYulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang et al.INFOCOM 2026
