COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM Accelerators
Taishu Sheng, Guangyu Sun, Dezun Dong
摘要
Chiplet-based architectures emerge as a promising approach to overcoming the physical and manufacturing constraints faced by monolithic chips, enabling the scalable integration of computing resources to meet the growing demands of AI workloads. However, efficient inter-chiplet communication still faces significant bottlenecks, especially under the fine-grained and bursty Direct Memory Access (DMA) request patterns generated by processing elements in modern AI tasks. Existing communication models and simulators fail to capture these characteristics, which limits the accuracy of performance analysis and the effectiveness of optimization strategies. These limitations hinder DMA-communication inefficiencies in chiplet-based AI systems and pose challenges for designing HPC architectures. To address these challenges, we present the first comprehensive chiplet communication model that explicitly incorporates finegrained DMA traffic observed in realistic AI workloads. Building on this model, we propose COMET, a novel framework that intelligently searches for optimal DMA request aggregation and memory address mapping strategies tailored to chiplet environments. COMET dynamically consolidates small DMA transfers to improve bandwidth utilization and reduce communication latency, while also adapting on-chip memory mapping to align with workload-specific dataflows. This mitigates synchronization overhead across diverse AI tasks. Compared with inference on conventional chiplet communication schemes, COMET achievesspeedup andhigher bandwidth utilization across different DNN and LLM workloads.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CHARM: Chiplet Heterogeneity-Aware Runtime Mapping SystemAlessandro Fogli, Bo Zhao, Peter R. Pietzuch, Jana GicevaEuroSys 2026 · 被引用 2 次
- SuperMesh: Energy-Efficient Collective Communications for AcceleratorsSabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung KimMICRO 2025 · 被引用 3 次
- Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsJingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei 等HPCA 2024 · 被引用 65 次
- NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip AcceleratorsZhanhong Tan, Hongyu Cai, Runpei Dong, Kaisheng MaISCA 2021 · 被引用 67 次
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu 等DAC 2024
