Lune

INFOCOM2026Top-tier venue

EdgeSpec: Distributed Speculative Decoding for Large Language Models at Edge

Yulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang, Haipeng Yao, Yuan He, Yunhao Liu

2026Year

Abstract

The growing demand for deploying large language models (LLMs) in applications requiring low latency and strong privacy has brought increasing attention to edge inference. While distributing computation across edge devices helps alleviate resource bottlenecks, the sequential and layer-wise nature of autoregressive decoding still limits real-time deployment, leading to device underutilization, increased latency, and workload imbalance. Existing methods rely on heavyweight drafts and static policies, which cannot adapt to device heterogeneity and bandwidth variation, leading to limited parallelism and elevated latency. To bridge this gap, we propose EdgeSpec, a hierarchical speculative decoding framework at edges. Instead of relying on last-layer speculation, EdgeSpec enables intermediate transformer layers to emit draft tokens, which are validated and used to predict the next token in parallel across devices via attention masking. The system combines offline entropy profiling with an online Kalman filter to adaptively select confident trigger layers based on runtime load and input complexity. To further improve efficiency, EdgeSpec jointly optimizes model partitioning, entropy thresholds, and candidate sizes through a two-stage strategy that integrates offline planning with online adaptation. Experiments on LLaMA2 showed that EdgeSpec reduced inference latency by up to 17.1× while preserving output quality comparable to standard decoding.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 73eb75eb-fa15-4eb3-944e-fa19d71cb2b9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines