A Tensor Compiler for Unified Machine Learning Prediction Serving
Supun Nakandala, Karla Saur, Gyeong-In Yu, Konstantinos Karanasos, Carlo Curino, Markus Weimer, Matteo Interlandi
Abstract
Machine Learning (ML) adoption in the enterprise requires simpler and more efficient software infrastructure-the bespoke solutions typical in large web companies are simply untenable. Model scoring, the process of obtaining predictions from a trained model over new data, is a primary contributor to infrastructure complexity and cost as models are trained once but used many times. In this paper we propose HUMMINGBIRD, a novel approach to model scoring, which compiles featurization operators and traditional ML models (e.g., decision trees) into a small set of tensor operations. This approach inherently reduces infrastructure complexity and directly leverages existing investments in Neural Network compilers and runtimes to generate efficient computations for both CPU and hardware accelerators. Our performance results are intriguing: despite replacing imperative computations (e.g., tree traversals) with tensor computation abstractions, HUMMINGBIRD is competitive and often outperforms hand-crafted kernels on micro-benchmarks on both CPU and GPU, while enabling seamless end-to-end acceleration of ML pipelines. We have released HUMMINGBIRD as open source. * The work was done while the author was at Microsoft.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d6c8763-d0f1-425c-b6c2-5b0a9172da98Cited by top-tier papers17
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen et al.SIGMOD 2022 · 50 citations
- Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionSize Zheng, Siyuan Chen, Peidi Song, Renze Chen et al.HPCA 2023 · 46 citations
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 45 citations
- Tensors: An abstraction for general data processingDimitrios Koutsoukos, Supun Nakandala, Konstantinos Karanasos, Karla Saur et al.VLDB 2021 · 38 citations
Related papers
- Treebeard: An Optimizing Compiler for Decision Tree Based ML InferenceAshwin Prasad, Sampath Rajendra, Kaushik Rajan, R. Govindarajan et al.MICRO 2022 · 8 citations
- InferDB: In-Database Machine Learning Inference Using IndexesRicardo Salazar-Díaz, Boris Glavic, Tilmann RablVLDB 2024 · 13 citations
- Lowering the Latency of Data Processing Pipelines Through FPGA based Hardware AccelerationMuhsen Owaida, Gustavo Alonso, Laura Fogliarini, Anthony Hock-Koon et al.VLDB 2020 · 48 citations
- Automatically Generating ML Compiler Backends from Tensor Accelerator ISA DescriptionsDevansh Jain, Akash Pardeshi, Marco Frigo, Kaustubh Khulbe et al.OOPSLA 2026
- RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding ColumnsZaifeng Pan, Zhen Zheng, Feng Zhang, Ruofan Wu et al.ASPLOS 2023 · 7 citations
