COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM Accelerators
Taishu Sheng, Guangyu Sun, Dezun Dong
Abstract
Chiplet-based architectures emerge as a promising approach to overcoming the physical and manufacturing constraints faced by monolithic chips, enabling the scalable integration of computing resources to meet the growing demands of AI workloads. However, efficient inter-chiplet communication still faces significant bottlenecks, especially under the fine-grained and bursty Direct Memory Access (DMA) request patterns generated by processing elements in modern AI tasks. Existing communication models and simulators fail to capture these characteristics, which limits the accuracy of performance analysis and the effectiveness of optimization strategies. These limitations hinder DMA-communication inefficiencies in chiplet-based AI systems and pose challenges for designing HPC architectures. To address these challenges, we present the first comprehensive chiplet communication model that explicitly incorporates finegrained DMA traffic observed in realistic AI workloads. Building on this model, we propose COMET, a novel framework that intelligently searches for optimal DMA request aggregation and memory address mapping strategies tailored to chiplet environments. COMET dynamically consolidates small DMA transfers to improve bandwidth utilization and reduce communication latency, while also adapting on-chip memory mapping to align with workload-specific dataflows. This mitigates synchronization overhead across diverse AI tasks. Compared with inference on conventional chiplet communication schemes, COMET achievesspeedup andhigher bandwidth utilization across different DNN and LLM workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 98e04302-4223-4b74-b37c-7734c8e61863Related papers
- CHARM: Chiplet Heterogeneity-Aware Runtime Mapping SystemAlessandro Fogli, Bo Zhao, Peter R. Pietzuch, Jana GicevaEuroSys 2026 · 2 citations
- SuperMesh: Energy-Efficient Collective Communications for AcceleratorsSabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung KimMICRO 2025 · 3 citations
- Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsJingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei et al.HPCA 2024 · 65 citations
- NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip AcceleratorsZhanhong Tan, Hongyu Cai, Runpei Dong, Kaisheng MaISCA 2021 · 67 citations
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu et al.DAC 2024
