MonoNN: Enabling a New Monolithic Optimization Space for Neural Network Inference Tasks on Modern GPU-Centric Architectures
Donglin Zhuang, Zhen Zheng, Haojun Xia, Xiafei Qiu, Junjie Bai, Wei Lin, Shuaiwen Leon Song
Abstract
In this work, we reveal that the kernel-by-kernel execution scheme in the existing machine learning optimizing compilers is no longer effective in fully utilizing hardware resources provided by the advances of modern GPU architectures. Specifically, such scheme suffers from severe non-computation overhead and off-chip memory traffic, making the optimization efforts from the state-of-the-art compiler techniques greatly attenuated on the newer generations of GPUs. To address this emerging challenge, we propose MonoNN, the first machine learning optimizing compiler that enables a new monolithic design and optimization space for common static neural network (NN) inference tasks on a single GPU. MonoNN can accommodate an entire neural network into a single GPU kernel, drastically reducing non-computation overhead and providing further fine-grained optimization opportunities from the newly formed monolithic optimization space. Most importantly, MonoNN identifies the resource incompatibility issue between various NN operators as the key design bottleneck for creating such a monolithic optimization space. Then MonoNN effectively tackles it by systematically exploring and exploiting the parallelism compensation strategy and resource trade-offs across different types of NN computations, and by proposing a novel scheduleindependent group tuning technique to significantly shrink the extremely large tuning space. Finally, MonoNN provides a compiler implementation that incorporates our proposed optimizations and automatically generates highly efficient kernel code. Extensive evaluation on a set of popular production inference tasks demonstrates that MonoNN achieves an average speedup of 2.01× over the state-of-the-art frameworks and compilers. Specifically, MonoNN outperforms TVM, TensorRT, XLA, and AStitch by up to 7.3×, 5.9×, 1.7× and 2.9× in terms of end-to-end inference performance, respectively. MonoNN source code is publicly available at https://github.com/AlibabaResearch/mononn.
⋄ Work was done when interned at Alibaba.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7739250-aa80-44c5-bc85-5914b086c5deCited by top-tier papers6
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective PrimitiveXinhao Luo, Zihan Liu, Yangjie Zhou, Shihan Fang et al.NeurIPS 2025 · 9 citations
- Characterizing Mobile SoC for Accelerating Heterogeneous LLM InferenceLe Chen, Dahu Feng, Erhu Feng, Yingrui Wang et al.SOSP 2025 · 3 citations
- RecFlex: Enabling Feature Heterogeneity-Aware Optimization for Deep Recommendation Models with Flexible SchedulesZaifeng Pan, Zhen Zheng, Feng Zhang, Bing Xie et al.SC 2024 · 2 citations
- MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM ServingDimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee et al.SC 2025 · 2 citations
- FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core ConnectionZiyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo et al.HPCA 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal et al.PLDI 2021 · 166 citations
Related papers
- RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding ColumnsZaifeng Pan, Zhen Zheng, Feng Zhang, Ruofan Wu et al.ASPLOS 2023 · 7 citations
- AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesZhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long et al.ASPLOS 2022 · 78 citations
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet et al.ASPLOS 2022 · 68 citations
- GraCE: Unlocking CUDA Graphs with Compiler Support for ML WorkloadsAbhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava BasuOSDI 2026
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 27 citations
