FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge Devices
Xiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao, Yunxin Liu
摘要
Due to the popularity of deep neural networks (DNNs) and considerations over network overhead, data privacy, and inference latency, there is a growing interest in deploying DNNs to edge devices in recent years. However, the limited memory becomes a major bottleneck for on-device DNN deployment, making it crucial to reduce the memory footprint of DNN. The mainstream model customization solutions require intensive deployment efforts and may lead to severe accuracy degradation, and existing deep learning (DL) frameworks don't take memory as a priority. Besides, recent works to enhance the memory management scheme cannot be directly applied because of several challenges, including the unbalanced memory footprint across layers, the inevitable overhead of memory management, and the memory budget dynamicity. To tackle these challenges, we introduce FlexNN, an efficient and adaptive memory management framework for DNN inference on memory-constrained devices. FlexNN uses a slicing-loading-computing joint planning approach, to achieve optimal memory utilization and minimal memory management overhead. We implemented FlexNN atop NCNN, and conducted comprehensive evaluations with common model architectures on various devices. The results have shown that our approach is able to adapt to different memory constraints with optimal latency-memory trade-offs. For example, FlexNN can reduce the memory consumption by 93.81% with only a 3.64% increase in latency, as compared with the original NCNN on smartphones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuningQitao Tan, Jun Liu, Zheng Zhan, Caiwen Ding 等NeurIPS 2025 · 被引用 19 次
- D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM ServingHaodong Wang, Qihua Zhou, Zicong Hong, Song GuoMobiCom 2025 · 被引用 8 次
- Nova: Real-Time Agentic Vision-Language Model Serving With Adaptive Cross-Stage ParallelizationYuhang Xu, Shengzhong Liu, Dong Zhang, Bingheng Yan 等RTSS 2025 · 被引用 1 次
- Learning Systems Expansion with Efficient Heterogeneity-aware Knowledge TransferGaole Dai, Huatao Xu, Yifan Yang, Rui Tan 等AAAI 2026
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu 等ASPLOS 2026
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li 等ICML 2023 · 被引用 683 次
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
相关 Paper
- BEEMS: Boosting Machine Vision Efficiency via Computation Graph-Based Memory SmoothingHanjing Shen, Fangxin Liu, Jian Liu, Li Jiang 等PPoPP 2026
- MASI: Memory-Adaptive Inference Framework for Spiking Neural Networks on Edge DevicesDi Yu, Helin Zheng, Changze Lv, Xin Du 等WWW 2026
- MoteNN: Memory Optimization via Fine-grained Scheduling for Deep Neural Networks on Tiny DevicesRenze Chen, Zijian Ding, Size Zheng, Meng Li 等DAC 2024 · 被引用 7 次
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato 等INFOCOM 2024 · 被引用 15 次
- Towards Memory-Efficient Neural Networks via Multi-Level in situ GenerationJiaqi Gu, Hanqing Zhu, Chenghao Feng, Mingjie Liu 等ICCV 2021 · 被引用 4 次
