Demand Layering for Real-Time DNN Inference with Minimized Memory Usage
Mingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn, Dongjoo Seo, Nikil D. Dutt, Jong-Chan Kim
摘要
When executing a deep neural network (DNN), its model parameters are loaded into GPU memory before execution, incurring a significant GPU memory burden. There are studies that reduce GPU memory usage by exploiting CPU memory as a swap device. However, this approach is not applicable in most embedded systems with integrated GPUs where CPU and GPU share a common memory. In this regard, we present Demand Layering, which employs a fast solid-state drive (SSD) as a co-running partner of a GPU and exploits the layer-by-layer execution of DNNs. In our approach, a DNN is loaded and executed in a layer-by-layer manner, minimizing the memory usage to the order of a single layer. Also, we developed a pipeline architecture that hides most additional delays caused by the interleaved parameter loadings alongside layer executions. Our implementation shows a 96.5% memory reduction with just 14.8% delay overhead on average for representative DNNs. Furthermore, by exploiting the memory-delay tradeoff, near-zero delay overhead (under 1 ms) can be achieved with a slightly increased memory usage (still an 88.4% reduction), showing the great potential of Demand Layering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsYufei Li, Zexin Li, Wei Yang, Cong LiuRTSS 2023 · 被引用 10 次
- : On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsZexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio 等RTSS 2023 · 被引用 9 次
- RT-MDM: Real-Time Scheduling Framework for Multi-DNN on MCU Using External MemorySukmin Kang, Seongtae Lee, Hyunwoo Koo, Hoon Sung Chwa 等DAC 2024 · 被引用 1 次
- Nova: Real-Time Agentic Vision-Language Model Serving With Adaptive Cross-Stage ParallelizationYuhang Xu, Shengzhong Liu, Dong Zhang, Bingheng Yan 等RTSS 2025 · 被引用 1 次
它引用的顶会 Paper12
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn 等NeurIPS 2020 · 被引用 827 次
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen 等ICML 2021 · 被引用 283 次
- Memory-efficient Patch-based Inference for Tiny Deep LearningJi Lin, Wei-Ming Chen, Han Cai, Chuang Gan 等NeurIPS 2021 · 被引用 190 次
- HetPipe: Enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data ParallelismJay H. Park, Gyeongchan Yun, Chang M. Yi, Nguyen T. Nguyen 等USENIX ATC 2020 · 被引用 178 次
- SwapAdvisor: Pushing Deep Learning Beyond the GPU Memory Limit via Smart SwappingChien-Chin Huang, Gu Jin, Jinyang LiASPLOS 2020 · 被引用 161 次
相关 Paper
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu 等ASPLOS 2026
- ZnG: Architecting GPU Multi-Processors with New Flash for Scalable Data AnalysisJie Zhang, Myoungsoo JungISCA 2020 · 被引用 14 次
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang 等DAC 2021 · 被引用 13 次
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin 等RTSS 2021 · 被引用 68 次
- Harmony: Overcoming the hurdles of GPU memory capacity to train massive DNN models on commodity serversYoujie Li, Amar Phanishayee, Derek Murray, Jakub Tarnawski 等VLDB 2022 · 被引用 28 次
