Transferable Graph Optimizers for ML Compilers
Yanqi Zhou, Sudip Roy, AmirAli Abdolrashidi, Daniel Wong, Peter C. Ma, Qiumin Xu, Hanxiao Liu, Mangpo Phitchaya Phothilimtha, Shen Wang, Anna Goldie, Azalia Mirhoseini, James Laudon
摘要
Most compilers for machine learning (ML) frameworks need to solve many correlated optimization problems to generate efficient machine code. Current ML compilers rely on heuristics based algorithms to solve these optimization problems one at a time. However, this approach is not only hard to maintain but often leads to sub-optimal solutions especially for newer model architectures. Existing learning based approaches in the literature are sample inefficient, tackle a single optimization problem, and do not generalize to unseen graphs making them infeasible to be deployed in practice. To address these limitations, we propose an end-to-end, transferable deep reinforcement learning method for computational graph optimization (GO), based on a scalable sequential attention mechanism over an inductive graph neural network. GO generates decisions on the entire graph rather than on each individual node autoregressively, drastically speeding up the search compared to prior methods. Moreover, we propose recurrent attention layers to jointly optimize dependent graph optimization tasks and demonstrate 33%-60% speedup on three graph optimization tasks compared to TensorFlow default optimization. On a diverse set of representative graphs consisting of up to 80,000 nodes, including Inception-v3, Transformer-XL, and WaveNet, GO achieves on average 21% improvement over human experts and 18% improvement over the prior state of the art with 15x faster convergence, on a device placement task evaluated in real systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Piper: Multidimensional Planner for DNN ParallelizationJakub Tarnawski, Deepak Narayanan, Amar PhanishayeeNeurIPS 2021 · 被引用 82 次
- HASCO: Towards Agile HArdware and Software CO-design for Tensor ComputationQingcheng Xiao, Size Zheng, Bingzhe Wu, Pengcheng Xu 等ISCA 2021 · 被引用 73 次
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu 等ASPLOS 2022 · 被引用 48 次
- LISA: Graph Neural Network based Portable Mapping on Spatial AcceleratorsZhaoying Li, Dan Wu, Dhananjaya Wijerathne, Tulika MitraHPCA 2022 · 被引用 43 次
- Learning Compiler Pass Orders using Coreset and Normalized Value PredictionYouwei Liang, Kevin Stone, Ali Shameli, Chris Cummins 等ICML 2023 · 被引用 26 次
相关 Paper
- Reinforced Genetic Algorithm Learning for Optimizing Computation GraphsAditya Paliwal, Felix Gimeno, Vinod Nair, Yujia Li 等ICLR 2020 · 被引用 70 次
- Neural Topological Ordering for Computation GraphsMukul Gagrani, Corrado Rainone, Yang Yang, Harris Teague 等NeurIPS 2022 · 被引用 21 次
- A Structure-Aware Framework for Learning Device Placements on Computation GraphsShukai Duan, Heng Ping, Nikos Kanakaris, Xiongye Xiao 等NeurIPS 2024 · 被引用 19 次
- RESPECT: Reinforcement Learning based Edge Scheduling on Pipelined Coral Edge TPUsJiaqi Yin, Yingjie Li, Daniel Robinson, Cunxi YuDAC 2023 · 被引用 9 次
- PluS: Highly Efficient and Expandable ML Compiler with Pluggable Graph SchedulesRuofan Wu, Zhen Zheng, Feng Zhang, Chuanjie Liu 等USENIX ATC 2025 · 被引用 5 次
