NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip Accelerators
Zhanhong Tan, Hongyu Cai, Runpei Dong, Kaisheng Ma
Abstract
The revolution of machine learning poses an unprecedented demand for computation resources, urging more transistors on a single monolithic chip, which is not sustainable in the Post-Moore era. The multichip integration with small functional dies, called chiplets, can reduce the manufacturing cost, improve the fabrication yield, and achieve die-level reuse for different system scales. DNN workload mapping and hardware design space exploration on such multichip systems are critical, but missing in the current stage.
This work provides a hierarchical and analytical framework to describe the DNN mapping on a multichip accelerator and analyze the communication overhead. Based on this framework, we propose an automatic tool called NN-Baton with a pre-design flow and a post-design flow. The pre-design flow aims to guide the chiplet granularity exploration with given area and performance budgets for the target workload. The post-design flow focuses on the workload orchestration on different computation levelspackage, chiplet, and core -in the hierarchy. Compared to Simba, NN-Baton generates mapping strategies that save 22.5%∼44% energy under the same computation and memory configurations. The architecture exploration demonstrates that area is a decisive factor for the chiplet granularity. For a 2048-MAC system under a 2 mm 2 chiplet area constraint, the 4-chiplet implementation with 4 cores and 16 lanes of 8-size vector-MAC is always the top-pick computation allocation across several benchmarks. In contrast, the optimal memory allocation policy in the hierarchy typically depends on the neural network models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 547436fb-9233-4e4c-a8de-44a27a214e43Cited by top-tier papers9
- Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsJingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei et al.HPCA 2024 · 65 citations
- SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module AcceleratorsMohanad Odema, Luke Chen, Hyoukjun Kwon, Mohammad Abdullah Al FaruqueMICRO 2024 · 11 citations
- Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication OptimizationZhanhong Tan, Zijian Zhu, Kaisheng MaASPLOS 2024 · 9 citations
- SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN AcceleratorsJingwei Cai, Xuan Wang, Mingyu Gao, Sen Peng et al.HPCA 2025 · 6 citations
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin et al.DAC 2023 · 5 citations
Builds on2
- Interstellar: Using Halide's Scheduling Language to Analyze DNN AcceleratorsXuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter et al.ASPLOS 2020 · 237 citations
- Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized RecommendationsRanggi Hwang, Taehun Kim, Youngeun Kwon, Minsoo RhuISCA 2020 · 94 citations
Related papers
- COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM AcceleratorsTaishu Sheng, Guangyu Sun, Dezun DongHPCA 2026
- On Design Space Exploration of Cache System in Multi-Chiplet SystemsYan Zhang, Xiaohang Wang, Yingtao Jiang, Amit Kumar SinghDAC 2025
- SPACX: Silicon Photonics-based Scalable Chiplet Accelerator for DNN InferenceYuan Li, Ahmed Louri, Avinash KaranthHPCA 2022 · 32 citations
- Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoCSize Zheng, Siyuan Chen, Yun LiangDAC 2023 · 10 citations
- Atomic Dataflow based Graph-Level Workload Orchestration for Scalable DNN AcceleratorsShixuan Zheng, Xianjue Zhang, Leibo Liu, Shaojun Wei et al.HPCA 2022 · 39 citations
