CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Qingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang, Dong Du, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen
摘要
Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Fully-Connected (FC) operations on GPUs. In this paper, we first design a Disaggregated Roofline Model (DRM) to characterize AFD performance, revealing that system throughput is constrained by the accelerator's limiting factor: either memory bandwidth or capacity. We observe that prior AFD systems often overlook these constraints and fail to balance them, leading to resource underutilization or constrained throughput. Therefore, we propose CHIME, the first AFD system integrating DIMM-PIM, which is a case of the new accelerator that strikes the balance with scalable capacity and bandwidth. To address the synchronization challenges inherent to the distributed cooperating DRAM chips in DIMM-PIM, CHIME employs bubble-free pipelining and hybrid-grained re-layout for efficient attention computation. Furthermore, it maximizes cross-device resource utilization via rankset-granular communication-computation overlapping and alignment-predicting scheduling. Evaluations show CHIME achieves up to speedup over state-of-theart HBM-PIM solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li 等ICML 2023 · 被引用 683 次
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan 等OSDI 2024 · 被引用 537 次
相关 Paper
- CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated MemoryXuchuan Luo, Jiacheng Shen, Pengfei Zuo, Xin Wang 等SOSP 2024 · 被引用 10 次
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU MultiplexingLei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong 等ICML 2026 · 被引用 4 次
- BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM InferenceZhichun Li, Jun Zhou, Xueqi Li, Ninghui SunDAC 2025 · 被引用 3 次
- POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM InferenceAditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 等ASPLOS 2025 · 被引用 24 次
- Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit ComputationHongyi Li, Songchen Ma, Huanyu Qu, Weihao Zhang 等ISCA 2026
