COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices
Yilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao, Jian Liu, Haibing Guan, Li Jiang
摘要
The development of on-device large language models (LLMs) is driven by the need for privacy and fast response times. Energyintensive data transfer on mobile devices makes Processing-in-Memory (PIM) an effective solution. Due to stringent DRAM cost constraints, limited physical footprint on circuit boards, and the interaction between applications and LLMs, it is imperative for the CPU and PIM to operate concurrently within a shared memory space. However, challenges such as bank conflicts and bus congestion can arise, potentially diminishing the performance and energy benefits of PIM.
To address this challenge, we introduce COSM, a cooperative scheduling framework designed to facilitate the concurrent operation of PIM and CPU tasks on mobile platforms. Our key innovations include: 1) a low-interference PIM control interface that generates the maximum number of PIM commands without disrupting CPU memory accesses; 2) an idleness-aware scheduling method that integrates PIM commands into available idle time windows within the CPU's access sequence. COSM not only hides PIM execution latency from the CPU, but also overlaps PIM execution with data transfer. Experiments on concurrent execution of LLMs and mobile workloads, including mobile applications and compute-intensive kernels, demonstrate that COSM improves PIM throughput by up to 2.8× compared to the baseline scheduling method with less than 2.0% CPU performance loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper34
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira 等ASPLOS 2021 · 被引用 182 次
- AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model InferenceJaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim 等ASPLOS 2024 · 被引用 125 次
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi 等ASPLOS 2024 · 被引用 121 次
- SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory SystemsMaciej Besta, Raghavendra Kanakagiri, Grzegorz Kwasniewski, Rachata Ausavarungnirun 等MICRO 2021 · 被引用 78 次
- GenStore: a high-performance in-storage processing system for genome sequence analysisNika Mansouri-Ghiasi, Jisung Park, Harun Mustafa, Jeremie S. Kim 等ASPLOS 2022 · 被引用 74 次
相关 Paper
- FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM InferenceSeong Hoon Seo, Junghoon Kim, Donghyun Lee, Seonah Yoo 等HPCA 2025 · 被引用 7 次
- PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-Based Long-Context LLM Inference SystemHyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee 等HPCA 2026 · 被引用 3 次
- COSPlay: Leveraging Task-Level Parallelism for High-Throughput Synchronous PersistenceMarina Vemmou, Alexandros DaglisMICRO 2021 · 被引用 1 次
- Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit ComputationHongyi Li, Songchen Ma, Huanyu Qu, Weihao Zhang 等ISCA 2026
- CoSine: Enhancing LLM Serving via Collaborative and Decoupled Speculative InferenceLuyao Gao, Jianchun Liu, Xichong Zhang, Guoju Gao 等INFOCOM 2026 · 被引用 1 次
