MAGIS: Memory Optimization via Coordinated Graph Transformation and Scheduling for DNN
Renze Chen, Zijian Ding, Size Zheng, Chengrui Zhang, Jingwen Leng, Xuanzhe Liu, Yun Liang
摘要
Recently, memory consumption of Deep Neural Network (DNN) rapidly increases, mainly due to long lifetimes and large shapes of tensors. Graph scheduling has emerged as an effective memory optimization technique, which determines the optimal execution, re-computation, swap-out, and swap-in timings for each operator/tensor. However, it often hurts performance significantly and can only manipulate tensors' lifetimes but not shapes, limiting the optimization space. We find that graph transformation, which can change the tensor shapes and graph structure, creates a new tradeoff space between memory and performance. Nevertheless, graph transformation are applied separately so far, with primary focus on optimizing performance and not memory.
In this paper, we propose MAGIS, a DNN memory optimization framework that coordinates graph transformation with graph scheduling. MAGIS uses a hierarchical tree to represent Fission Transformation (F-Trans), a type of transformation which can effectively reduce tensor shapes in a sub-graph. To keep the complexity low, we build a light-weight search space based on graph structure analysis. MAGIS decomposes graph scheduling into graph transformation and re-ordering and designs an incremental scheduling * Work done while the author was a student at Peking University.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu 等NeurIPS 2024 · 被引用 56 次
- FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core ConnectionZiyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo 等HPCA 2026
- MTP: A Meaning-Typed Language Abstraction for AI-Integrated ProgrammingJayanaka L. Dantanarayana, Yiping Kang, Kugesan Sivasothynathan, Christopher Clarke 等OOPSLA 2025
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
相关 Paper
- Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication OptimizationZhanhong Tan, Zijian Zhu, Kaisheng MaASPLOS 2024 · 被引用 9 次
- MoteNN: Memory Optimization via Fine-grained Scheduling for Deep Neural Networks on Tiny DevicesRenze Chen, Zijian Ding, Size Zheng, Meng Li 等DAC 2024 · 被引用 7 次
- Piper: Multidimensional Planner for DNN ParallelizationJakub Tarnawski, Deepak Narayanan, Amar PhanishayeeNeurIPS 2021 · 被引用 82 次
- EDA: Energy-Efficient Inter-Layer Model Compilation for Edge DNN Inference AccelerationBo Ren Pao, I-Chia Chen, En-Hao Chang, Tsung Tai YehHPCA 2025 · 被引用 1 次
- SFD: Towards Segment Fusion Dataflow for Spatial AcceleratorsFuyu Wang, Minghua Shen, Yufei Ding, Nong Xiao 等HPCA 2026
