SHyLA: 3D-Stacked NVM-DRAM Hybrid LLM-Inference Architecture Exploiting Data and Memory Heterogeneity
Liu He, Fuyao Zhou, Cheng Peng, Shunan Dong, Ziming Zhang, Huazhong Yang, Yongpan Liu, Guowei Zhang, Hongyang Jia
Abstract
Generative large language models (LLMs) in cloud services impose substantial memory demands due to massive parameters and large key-value caches, especially in long-context scenarios. To overcome bandwidth and capacity limits of conventional memory systems, hybrid memory and 3D-stacked architectures have emerged as promising solutions. However, prior Non-Volatile Memory (NVM)-DRAM hybrid architecture studies often characterize NVM as dense but bandwidth-limited storage, overlooking its potential for favorable bandwidth-capacity tradeoffs enabled by the areal bandwidth scaling in true 3D stacking. To tackle this problem, our work jointly considers memory and LLM data heterogeneity to reveal opportunities at the workload-hardware interface. Guided by this characterization, we propose SHyLA, a hardware-software heterogeneity-aware 3D-stacked NVM-DRAM hybrid architecture for LLM inference. SHyLA strategically places different LLM data categories across memory devices and employs a bandwidth-utilization-centric dataflow that exploits the areal bandwidth of 3D stacking, particularly under hybrid memory data placement constraints. A two-stage design space exploration methodology further navigates the expanded hybrid memory and deployment design space to maximize system throughput under per-user throughput constraints with architectural insights. Evaluations demonstrate that SHyLA achieves up to geomean) over a DRAM-only baseline and up to (1.76× geomean) over an NVM-only baseline, while maintaining acceptable thermal behavior and a practical lifetime.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 015f1784-e4a5-430a-bf91-14e10f30d3b5Related papers
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE ServingYue Pan, Zihan Xia, Po-Kai Hsu, Lanxiang Hu et al.MICRO 2025 · 7 citations
- Lincoln: Real-Time 50 100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash MemoryWeiyi Sun, Mingyu Gao, Zhaoshi Li, Aoyang Zhang et al.HPCA 2025 · 11 citations
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu et al.ISCA 2026 · 4 citations
- HybridSpec: Exploiting Hybrid-Bonding Memory to Accelerate LLM Serving Through Heterogeneous Architecture and Speculative DecodingZongle Huang, Wenbin Jia, Yaolei Li, Xinyuan Lin et al.ISCA 2026 · 2 citations
- Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid BondingSanghyeok Han, Byungkuk Yoon, Gyeonghwan Park, Choungki Song et al.DAC 2025 · 4 citations
