Lune

ISCA2026顶会

SHyLA: 3D-Stacked NVM-DRAM Hybrid LLM-Inference Architecture Exploiting Data and Memory Heterogeneity

Liu He, Fuyao Zhou, Cheng Peng, Shunan Dong, Ziming Zhang, Huazhong Yang, Yongpan Liu, Guowei Zhang, Hongyang Jia

2026年份
1被引次数

摘要

Generative large language models (LLMs) in cloud services impose substantial memory demands due to massive parameters and large key-value caches, especially in long-context scenarios. To overcome bandwidth and capacity limits of conventional memory systems, hybrid memory and 3D-stacked architectures have emerged as promising solutions. However, prior Non-Volatile Memory (NVM)-DRAM hybrid architecture studies often characterize NVM as dense but bandwidth-limited storage, overlooking its potential for favorable bandwidth-capacity tradeoffs enabled by the areal bandwidth scaling in true 3D stacking. To tackle this problem, our work jointly considers memory and LLM data heterogeneity to reveal opportunities at the workload-hardware interface. Guided by this characterization, we propose SHyLA, a hardware-software heterogeneity-aware 3D-stacked NVM-DRAM hybrid architecture for LLM inference. SHyLA strategically places different LLM data categories across memory devices and employs a bandwidth-utilization-centric dataflow that exploits the areal bandwidth of 3D stacking, particularly under hybrid memory data placement constraints. A two-stage design space exploration methodology further navigates the expanded hybrid memory and deployment design space to maximize system throughput under per-user throughput constraints with architectural insights. Evaluations demonstrate that SHyLA achieves up to 5.84×(2.02×5.84 \times (\mathbf{2. 0 2} \times geomean) over a DRAM-only baseline and up to 6.03×\mathbf{6. 0 3} \times (1.76× geomean) over an NVM-only baseline, while maintaining acceptable thermal behavior and a practical lifetime.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖