PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory
Shuang Chen, Yi Jiang, Christina Delimitrou, José F. Martínez
Abstract
The slowdown of Moore’s Law, combined with advances in 3D stacking of logic and memory, have pushed architects to revisit the concept of processing-in-memory (PIM) to overcome the memory wall bottleneck. This PIM renaissance finds itself in a very different computing landscape from the one twenty years ago, as more and more computation shifts to the cloud. Most PIM architecture papers still focus on best-effort applications, while PIM’s impact on latency-critical cloud applications is not well understood.This paper explores how datacenters can exploit PIM architectures in the context of latency-critical applications. We adopt a general-purpose cloud server with HBM-based, 3D-stacked logic+memory modules, and study the impact of PIM on six diverse interactive cloud applications. We reveal the previously neglected opportunity that PIM presents to these services, and show the importance of properly managing PIM-related resources to meet the QoS targets of interactive services and maximize resource efficiency. Then, we present PIMCloud, a QoS-aware resource manager designed for cloud systems with PIM allowing colocation of multiple latency-critical and best-effort applications. We show that PIMCloud efficiently manages PIM resources: it (1) improves effective machine utilization by up to 70% and 85% (average 24% and 33%) under 2-app and 3-app mixes, compared to the best state-of-the-art manager; (2) helps latency-critical applications meet QoS; and (3) adapts to varying load patterns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 735d2a03-7cdc-4f86-b6e6-a2326b4442ddCited by top-tier papers4
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 62 citations
- ABNDP: Co-optimizing Data Access and Load Balance in Near-Data ProcessingBoyu Tian, Qihang Chen, Mingyu GaoASPLOS 2023 · 31 citations
- PIM-Malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) ArchitecturesDongjae Lee, Bongjoon Hyun, Youngjin Kwon, Minsoo RhuHPCA 2026 · 1 citation
Builds on2
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 153 citations
- Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud ServicesRajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus SjälanderHPCA 2020 · 76 citations
Related papers
- vPIM: Efficient Virtual Address Translation for Scalable Processing-in-Memory ArchitecturesAmel Fatima, Sihang Liu, Korakit Seemakhupt, Rachata Ausavarungnirun et al.DAC 2023 · 6 citations
- A Case Study of Processing-in-Memory in off-the-Shelf SystemsJoel Nider, Craig Mustard, Andrada Zoltan, John Ramsden et al.USENIX ATC 2021 · 62 citations
- PIM-CCA: An Efficient PIM Architecture with Optimized Integration of Configurable Functional UnitsJeehyun Kim, Donghyeon Kim, Seokwon Kang, Bongjoon Hyun et al.MICRO 2025 · 3 citations
- PIM-tree: A Skew-resistant Index for Processing-in-MemoryHongbo Kang, Yiwei Zhao, Guy E. Blelloch, Laxman Dhulipala et al.VLDB 2023 · 40 citations
- HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI DevicesSangmin Jeon, Kangju Lee, Kyeongwon Lee, Woojoo LeeDAC 2025 · 4 citations
