Early Silicon of Raptor: The First 3D-DRAM Accelerator for Generative Inference
Prashant J. Nair, Ramyad Hadidi, Subramani Ganesh, Sangamesh Kodge, Shubhankit Rathore, Neil Thanawala, Nikitha Reddy, Gyanesh Saharia, Vinayak Patankar, Arun Tiruvur, Nithesh Kurella, Sudeep Bhoja
摘要
Generative Inference is largely memory-bound. Autoregressive decoding dominates inference runtime and thus drives memory bandwidth and capacity demands. SRAM-based accelerators offer high bandwidth but limited capacity, while HBM DRAM provides capacity but is constrained by bandwidth and power. 3D-stacked logic-on-DRAM (3D-DRAM) helps close this gap, but integrating 3D-DRAM with accelerator logic introduces four challenges: (1) workload-aware mapping to exploit parallelism, (2) power optimization without burst-based data bus inversion (DBI), (3) resilience with high bank counts, and (4) thermal reliability at elevated junction temperatures. This paper distills lessons from the early silicon of Raptor, the first commercial generative inference accelerator with 3Dstacked DRAM (3D-DRAM). This paper introduces four key architectural features to enable practical 3D-DRAM integration. First, stream-blocking maps KV-cache streams onto configurable 3D-DRAM channels, sustaining up to 100 TB/s per card while preserving bank-level parallelism. Second, pinless DBI on the single-cycle bump interface reduces memory-subsystem energy. Third, topology-preserving redundancy with thermal-aware refresh. Fourth, interleaved ECC for reliability at high temperatures. Across Llama-3.1 70B, DeepSeek-V3, Kimi K2, GPT-OSS, Whisper, and Canary, Raptor 3D-DRAM gives and higher throughput than with HBM and SRAM, respectively, while being less sensitive to network latency and bandwidth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper40
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng 等ICML 2024 · 被引用 669 次
相关 Paper
- 3D-TokSIM: Stacking 3D Memory with Token-Stationary Compute-in-Memory for Speculative LLM InferenceWentao Zhao, Boya Lv, Meng Wu, Peiyu Chen 等DAC 2025 · 被引用 4 次
- SHyLA: 3D-Stacked NVM-DRAM Hybrid LLM-Inference Architecture Exploiting Data and Memory HeterogeneityLiu He, Fuyao Zhou, Cheng Peng, Shunan Dong 等ISCA 2026 · 被引用 1 次
- Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid BondingSanghyeok Han, Byungkuk Yoon, Gyeonghwan Park, Choungki Song 等DAC 2025 · 被引用 4 次
- HybridSpec: Exploiting Hybrid-Bonding Memory to Accelerate LLM Serving Through Heterogeneous Architecture and Speculative DecodingZongle Huang, Wenbin Jia, Yaolei Li, Xinyuan Lin 等ISCA 2026 · 被引用 2 次
- ATiM: Autotuning Tensor Programs for Processing-in-DRAMYongwon Shin, Dookyung Kang, Hyojin SungISCA 2025 · 被引用 2 次
