PluS: Highly Efficient and Expandable ML Compiler with Pluggable Graph Schedules
Ruofan Wu, Zhen Zheng, Feng Zhang, Chuanjie Liu, Zaifeng Pan, Jidong Zhai, Xiaoyong Du
Abstract
Machine learning (ML) compilers are effective solutions for deploying diverse Deep Neural Network (DNN) workloads on various hardware platforms automatically. However, there is a notable lag in existing ML compilers when it comes to supporting emerging optimization techniques like recent attention optimizations. These compilers lack the requisite flexibility to support expert-driven subgraph optimizations timely, resulting in suboptimal performance compared to manually optimized libraries. Conversely, template-based compilers lack the ability to abstractly express subgraphs, thereby reducing their adaptability to subtle changes in model architectures.
In this paper, we present PluS, an end-to-end ML compiler that facilitates the deployment of expert-optimized subgraph implementations while still preserving compiler flexibility. We rethink the encapsulation of ML compiler and decouple the burdensome embedded graph transformation process. PluS provides a lightweight loop-centric subgraph abstraction for experts to manage a flexible pattern warehouse, and employs a pattern identification approach for subgraph generation. As a result, PluS can deploy efficient subgraph implementations with minimal manual efforts, making it outperform the state-of-the-art rule-based embedded compilers (up to 4.04× speedup) on popular ML models.
- Work was done when Ruofan and Zaifeng interned at Microsoft, advised by Zhen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9e8bd1a-dfa7-43c4-82ca-bc856ba37248Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
Related papers
- Transferable Graph Optimizers for ML CompilersYanqi Zhou, Sudip Roy, AmirAli Abdolrashidi, Daniel Wong et al.NeurIPS 2020 · 63 citations
- RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding ColumnsZaifeng Pan, Zhen Zheng, Feng Zhang, Ruofan Wu et al.ASPLOS 2023 · 7 citations
- EVT: Accelerating Deep Learning Training with Epilogue Visitor TreeZhaodong Chen, Andrew Kerr, Richard Cai, Jack Kosaian et al.ASPLOS 2024 · 5 citations
- DyCL: Dynamic Neural Network Compilation Via Program Rewriting and Graph OptimizationSimin Chen, Shiyi Wei, Cong Liu, Wei YangISSTA 2023 · 11 citations
- ALT: Breaking the Wall between Data Layout and Loop Optimizations for Deep Learning CompilationZhiying Xu, Jiafan Xu, Hongding Peng, Wei Wang et al.EuroSys 2023 · 12 citations
