HybridSpec: Exploiting Hybrid-Bonding Memory to Accelerate LLM Serving Through Heterogeneous Architecture and Speculative Decoding
Zongle Huang, Wenbin Jia, Yaolei Li, Xinyuan Lin, Shupei Fan, Chen Tang, Shuwen Deng, Yongpan Liu
摘要
Large Language Models (LLMs) are widely used in latency-sensitive serving. Their autoregressive inference demands both high memory bandwidth for token generation and large capacity for massive parameters and growing KV caches. This dual requirement poses a fundamental challenge for memory systems. Hybrid bonding (HB), a recent packaging advancement that delivers unprecedented bandwidth via 3D DRAM-logic stacking, yet faces capacity limitations that confine prior designs to small models or edge settings. We identify that speculative decoding (SD), a lossless method where a lightweight draft model proposes draft tokens that are then verified in parallel by the target model, polarizes memory demands and thus enables separate memory optimization under physical constraints. The low-arithmetic-intensity draft model requires high bandwidth but little capacity, while the computeintensive target model requires large capacity but tolerates lower bandwidth. Based on this insight, we present HybridSpec, a heterogeneous system that combines a high-bandwidth HB stack with large-capacity LPDDR5X. We adopt a mapping where communication between XPU and HB stack occurs only at draftverification boundaries, reducing data transfer. To fully exploit this architecture under dynamic online serving, we further propose: (1) asynchronous batching to balance latency and throughput; (2) utilization-aware speculation that adapts to runtime workloads; and (3) prefill-verification arbitration to mitigate resource contention. Evaluations show HybridSpec improves latency and energy by and over GPU baselines, supports higher request rates under the same servicelevel objectives (SLOs), and remains cost-effective. While prior heterogeneous designs emphasize arithmetic intensity matching, we reveal that absolute memory capacity can be a limiting factor, with significant impact on serving latency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative DecodingRanajoy Sadhukhan, Jian Chen, Zhuoming Chen, Vashisth Tiwari 等ICLR 2025
- Lincoln: Real-Time 50 100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash MemoryWeiyi Sun, Mingyu Gao, Zhaoshi Li, Aoyang Zhang 等HPCA 2025 · 被引用 11 次
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 被引用 20 次
- SwiftSpec: Disaggregated Speculative Decoding and Fused Kernels for Low-Latency LLM InferenceZiyi Zhang, Ziheng Jiang, Chengquan Jiang, Menghan Yu 等ASPLOS 2026
- AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative DecodingZikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro 等EuroSys 2026
