FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
Zhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu, Miao Yin, Gagan Agrawal, Wei Niu
摘要
The increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0× to 8.4× memory reduction and 1.7× to 75.0× speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang 等ASPLOS 2020 · 被引用 214 次
相关 Paper
- FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesXiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao 等MobiCom 2024 · 被引用 40 次
- Demand Layering for Real-Time DNN Inference with Minimized Memory UsageMingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn 等RTSS 2022 · 被引用 21 次
- SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on MobileWei Niu, Md. Musfiqur Rahman Sanim, Zhihao Shu, Jiexiong Guan 等ASPLOS 2024 · 被引用 11 次
- LLM in a flash: Efficient Large Language Model Inference with Limited MemoryKeivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, S. Khatamifard 等ACL 2024 · 被引用 73 次
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsIsmet Dagli, Mehmet E. BelviranliPPoPP 2024 · 被引用 18 次
