EdgeSpec: Distributed Speculative Decoding for Large Language Models at Edge
Yulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang, Haipeng Yao, Yuan He, Yunhao Liu
摘要
The growing demand for deploying large language models (LLMs) in applications requiring low latency and strong privacy has brought increasing attention to edge inference. While distributing computation across edge devices helps alleviate resource bottlenecks, the sequential and layer-wise nature of autoregressive decoding still limits real-time deployment, leading to device underutilization, increased latency, and workload imbalance. Existing methods rely on heavyweight drafts and static policies, which cannot adapt to device heterogeneity and bandwidth variation, leading to limited parallelism and elevated latency. To bridge this gap, we propose EdgeSpec, a hierarchical speculative decoding framework at edges. Instead of relying on last-layer speculation, EdgeSpec enables intermediate transformer layers to emit draft tokens, which are validated and used to predict the next token in parallel across devices via attention masking. The system combines offline entropy profiling with an online Kalman filter to adaptively select confident trigger layers based on runtime load and input complexity. To further improve efficiency, EdgeSpec jointly optimizes model partitioning, entropy thresholds, and candidate sizes through a two-stage strategy that integrates offline planning with online adaptation. Experiments on LLaMA2 showed that EdgeSpec reduced inference latency by up to 17.1× while preserving output quality comparable to standard decoding.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 被引用 20 次
- HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative DecodingSiran Liu, Yang Ye, Qianchao Zhu, Zane Cao 等ACL 2026 · 被引用 2 次
- QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV CacheRishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Richard Charles Hooper 等ICML 2025
- PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative DecodingHan Yunhe, Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi 等ICML 2026 · 被引用 2 次
- EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU UtilizationYize Wu, Ke Gao, Ling Li, Yanjun WuNeurIPS 2025 · 被引用 3 次
