Lune

INFOCOM2026顶会

EdgeSpec: Distributed Speculative Decoding for Large Language Models at Edge

Yulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang, Haipeng Yao, Yuan He, Yunhao Liu

2026年份

摘要

The growing demand for deploying large language models (LLMs) in applications requiring low latency and strong privacy has brought increasing attention to edge inference. While distributing computation across edge devices helps alleviate resource bottlenecks, the sequential and layer-wise nature of autoregressive decoding still limits real-time deployment, leading to device underutilization, increased latency, and workload imbalance. Existing methods rely on heavyweight drafts and static policies, which cannot adapt to device heterogeneity and bandwidth variation, leading to limited parallelism and elevated latency. To bridge this gap, we propose EdgeSpec, a hierarchical speculative decoding framework at edges. Instead of relying on last-layer speculation, EdgeSpec enables intermediate transformer layers to emit draft tokens, which are validated and used to predict the next token in parallel across devices via attention masking. The system combines offline entropy profiling with an online Kalman filter to adaptively select confident trigger layers based on runtime load and input complexity. To further improve efficiency, EdgeSpec jointly optimizes model partitioning, entropy thresholds, and candidate sizes through a two-stage strategy that integrates offline planning with online adaptation. Experiments on LLaMA2 showed that EdgeSpec reduced inference latency by up to 17.1× while preserving output quality comparable to standard decoding.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 73eb75eb-fa15-4eb3-944e-fa19d71cb2b9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖