EdgeSpec: Distributed Speculative Decoding for Large Language Models at Edge
Yulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang, Haipeng Yao, Yuan He, Yunhao Liu
Abstract
The growing demand for deploying large language models (LLMs) in applications requiring low latency and strong privacy has brought increasing attention to edge inference. While distributing computation across edge devices helps alleviate resource bottlenecks, the sequential and layer-wise nature of autoregressive decoding still limits real-time deployment, leading to device underutilization, increased latency, and workload imbalance. Existing methods rely on heavyweight drafts and static policies, which cannot adapt to device heterogeneity and bandwidth variation, leading to limited parallelism and elevated latency. To bridge this gap, we propose EdgeSpec, a hierarchical speculative decoding framework at edges. Instead of relying on last-layer speculation, EdgeSpec enables intermediate transformer layers to emit draft tokens, which are validated and used to predict the next token in parallel across devices via attention masking. The system combines offline entropy profiling with an online Kalman filter to adaptively select confident trigger layers based on runtime load and input complexity. To further improve efficiency, EdgeSpec jointly optimizes model partitioning, entropy thresholds, and candidate sizes through a two-stage strategy that integrates offline planning with online adaptation. Experiments on LLaMA2 showed that EdgeSpec reduced inference latency by up to 17.1× while preserving output quality comparable to standard decoding.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 73eb75eb-fa15-4eb3-944e-fa19d71cb2b9Related papers
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 20 citations
- HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative DecodingSiran Liu, Yang Ye, Qianchao Zhu, Zane Cao et al.ACL 2026 · 2 citations
- QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV CacheRishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Richard Charles Hooper et al.ICML 2025
- PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative DecodingHan Yunhe, Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi et al.ICML 2026 · 2 citations
- EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU UtilizationYize Wu, Ke Gao, Ling Li, Yanjun WuNeurIPS 2025 · 3 citations
