NutCracker: A Compilation Framework for Hybrid DPU Architectures
Yihan Yang, Haifeng Sun, Antoine Kaufmann, Jialin Li
Abstract
SoC-based SmartNICs, or data processing units (DPUs), are becoming a viable option for offloading infrastructure services. However, developers need deep hardware-level knowledge to fully utilize the available hardware accelerators on a target DPU. Hardware heterogeneity also makes porting across DPUs a formidable task. In this work, we propose a new compiler framework, NutCracker. Using NutCracker, programmers develop DPU applications using high-level target-independent languages. NutCracker applies a two-stage compilation process. It first performs progressive lowering to convert the source program to candidate intermediate representations (IRs) of the target DPU. Next, NutCracker applies cost-guided mapping optimization using equality saturation to select a final implementation on the target hardware with a configurable optimization goal. Evaluated on eight applications, NutCracker reduces developer effort by nearly 90% while delivering performance within 3% of handcrafted implementations for seven of the workloads. Moreover, its compilation time is comparable to standard toolchains such as GCC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e49ec4c1-8d87-47a1-8b5c-fd3e8d03ee7aCited by top-tier papers1
Ask how each one uses itBuilds on23
- egg: Fast and extensible equality saturationMax Willsey, Chandrakana Nandi, Yisu Remy Wang, Oliver Flatt et al.POPL 2021 · 170 citations
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im et al.SOSP 2021 · 83 citations
- Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsJiaqi Gao, Ennan Zhai, Hongqiang Harry Liu, Rui Miao et al.SIGCOMM 2020 · 82 citations
- FlexTOE: Flexible TCP Offload with Fine-Grained ParallelismRajath Shashidhara, Tim Stamler, Antoine Kaufmann, Simon PeterNSDI 2022 · 66 citations
- Xenic: SmartNIC-Accelerated Distributed TransactionsHenry N. Schuh, Weihao Liang, Ming Liu, Jacob Nelson et al.SOSP 2021 · 62 citations
Related papers
- Enabling Portable and High-Performance SmartNIC Programs with AlkaliJiaxin Lin, Zhiyuan Guo, Mihir Shah, Tao Ji et al.NSDI 2025 · 11 citations
- dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data ProcessingJiasheng Hu, Kaiwen Zheng, Anna Li, Sidharth Sankhe et al.VLDB 2026
- Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUsDimitrios Giouroukis, Dwi P. A. Nugroho, Varun Pandey, Steffen Zeuch et al.VLDB 2025 · 3 citations
- Conspirator: SmartNIC-Aided Control Plane for Distributed ML WorkloadsYunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao et al.USENIX ATC 2024 · 13 citations
- Taming the Zoo: The Unified GraphIt Compiler Framework for Novel ArchitecturesAjay Brahmakshatriya, Emily Furst, Victor A. Ying, Claire Hsu et al.ISCA 2021 · 12 citations
