Lune

ISCA2026顶会

HybridSpec: Exploiting Hybrid-Bonding Memory to Accelerate LLM Serving Through Heterogeneous Architecture and Speculative Decoding

Zongle Huang, Wenbin Jia, Yaolei Li, Xinyuan Lin, Shupei Fan, Chen Tang, Shuwen Deng, Yongpan Liu

2026年份
2被引次数

摘要

Large Language Models (LLMs) are widely used in latency-sensitive serving. Their autoregressive inference demands both high memory bandwidth for token generation and large capacity for massive parameters and growing KV caches. This dual requirement poses a fundamental challenge for memory systems. Hybrid bonding (HB), a recent packaging advancement that delivers unprecedented bandwidth via 3D DRAM-logic stacking, yet faces capacity limitations that confine prior designs to small models or edge settings. We identify that speculative decoding (SD), a lossless method where a lightweight draft model proposes draft tokens that are then verified in parallel by the target model, polarizes memory demands and thus enables separate memory optimization under physical constraints. The low-arithmetic-intensity draft model requires high bandwidth but little capacity, while the computeintensive target model requires large capacity but tolerates lower bandwidth. Based on this insight, we present HybridSpec, a heterogeneous system that combines a high-bandwidth HB stack with large-capacity LPDDR5X. We adopt a mapping where communication between XPU and HB stack occurs only at draftverification boundaries, reducing data transfer. To fully exploit this architecture under dynamic online serving, we further propose: (1) asynchronous batching to balance latency and throughput; (2) utilization-aware speculation that adapts to runtime workloads; and (3) prefill-verification arbitration to mitigate resource contention. Evaluations show HybridSpec improves latency and energy by 3.02×3.02 \times and 1.96×1.96 \times over GPU baselines, supports 2.97×2.97 \times higher request rates under the same servicelevel objectives (SLOs), and remains cost-effective. While prior heterogeneous designs emphasize arithmetic intensity matching, we reveal that absolute memory capacity can be a limiting factor, with significant impact on serving latency.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 57542b00-e1f7-4b06-9c9e-cd687772e29a

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖