Lune

ISCA2026Top-tier venue

HybridSpec: Exploiting Hybrid-Bonding Memory to Accelerate LLM Serving Through Heterogeneous Architecture and Speculative Decoding

Zongle Huang, Wenbin Jia, Yaolei Li, Xinyuan Lin, Shupei Fan, Chen Tang, Shuwen Deng, Yongpan Liu

2026Year
2Citations

Abstract

Large Language Models (LLMs) are widely used in latency-sensitive serving. Their autoregressive inference demands both high memory bandwidth for token generation and large capacity for massive parameters and growing KV caches. This dual requirement poses a fundamental challenge for memory systems. Hybrid bonding (HB), a recent packaging advancement that delivers unprecedented bandwidth via 3D DRAM-logic stacking, yet faces capacity limitations that confine prior designs to small models or edge settings. We identify that speculative decoding (SD), a lossless method where a lightweight draft model proposes draft tokens that are then verified in parallel by the target model, polarizes memory demands and thus enables separate memory optimization under physical constraints. The low-arithmetic-intensity draft model requires high bandwidth but little capacity, while the computeintensive target model requires large capacity but tolerates lower bandwidth. Based on this insight, we present HybridSpec, a heterogeneous system that combines a high-bandwidth HB stack with large-capacity LPDDR5X. We adopt a mapping where communication between XPU and HB stack occurs only at draftverification boundaries, reducing data transfer. To fully exploit this architecture under dynamic online serving, we further propose: (1) asynchronous batching to balance latency and throughput; (2) utilization-aware speculation that adapts to runtime workloads; and (3) prefill-verification arbitration to mitigate resource contention. Evaluations show HybridSpec improves latency and energy by 3.02×3.02 \times and 1.96×1.96 \times over GPU baselines, supports 2.97×2.97 \times higher request rates under the same servicelevel objectives (SLOs), and remains cost-effective. While prior heterogeneous designs emphasize arithmetic intensity matching, we reveal that absolute memory capacity can be a limiting factor, with significant impact on serving latency.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 57542b00-e1f7-4b06-9c9e-cd687772e29a

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines