Demand Layering for Real-Time DNN Inference with Minimized Memory Usage
Mingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn, Dongjoo Seo, Nikil D. Dutt, Jong-Chan Kim
Abstract
When executing a deep neural network (DNN), its model parameters are loaded into GPU memory before execution, incurring a significant GPU memory burden. There are studies that reduce GPU memory usage by exploiting CPU memory as a swap device. However, this approach is not applicable in most embedded systems with integrated GPUs where CPU and GPU share a common memory. In this regard, we present Demand Layering, which employs a fast solid-state drive (SSD) as a co-running partner of a GPU and exploits the layer-by-layer execution of DNNs. In our approach, a DNN is loaded and executed in a layer-by-layer manner, minimizing the memory usage to the order of a single layer. Also, we developed a pipeline architecture that hides most additional delays caused by the interleaved parameter loadings alongside layer executions. Our implementation shows a 96.5% memory reduction with just 14.8% delay overhead on average for representative DNNs. Furthermore, by exploiting the memory-delay tradeoff, near-zero delay overhead (under 1 ms) can be achieved with a slightly increased memory usage (still an 88.4% reduction), showing the great potential of Demand Layering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bda102d2-ba5b-4c36-8597-5a22ff4ce9c4Cited by top-tier papers4
- RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsYufei Li, Zexin Li, Wei Yang, Cong LiuRTSS 2023 · 10 citations
- : On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsZexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio et al.RTSS 2023 · 9 citations
- RT-MDM: Real-Time Scheduling Framework for Multi-DNN on MCU Using External MemorySukmin Kang, Seongtae Lee, Hyunwoo Koo, Hoon Sung Chwa et al.DAC 2024 · 1 citation
- Nova: Real-Time Agentic Vision-Language Model Serving With Adaptive Cross-Stage ParallelizationYuhang Xu, Shengzhong Liu, Dong Zhang, Bingheng Yan et al.RTSS 2025 · 1 citation
Builds on12
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen et al.ICML 2021 · 283 citations
- Memory-efficient Patch-based Inference for Tiny Deep LearningJi Lin, Wei-Ming Chen, Han Cai, Chuang Gan et al.NeurIPS 2021 · 190 citations
- HetPipe: Enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data ParallelismJay H. Park, Gyeongchan Yun, Chang M. Yi, Nguyen T. Nguyen et al.USENIX ATC 2020 · 178 citations
- SwapAdvisor: Pushing Deep Learning Beyond the GPU Memory Limit via Smart SwappingChien-Chin Huang, Gu Jin, Jinyang LiASPLOS 2020 · 161 citations
Related papers
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu et al.ASPLOS 2026
- ZnG: Architecting GPU Multi-Processors with New Flash for Scalable Data AnalysisJie Zhang, Myoungsoo JungISCA 2020 · 14 citations
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang et al.DAC 2021 · 13 citations
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin et al.RTSS 2021 · 68 citations
- Harmony: Overcoming the hurdles of GPU memory capacity to train massive DNN models on commodity serversYoujie Li, Amar Phanishayee, Derek Murray, Jakub Tarnawski et al.VLDB 2022 · 28 citations
