Lune

HPCA2026顶会

COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM Accelerators

Taishu Sheng, Guangyu Sun, Dezun Dong

2026年份

摘要

Chiplet-based architectures emerge as a promising approach to overcoming the physical and manufacturing constraints faced by monolithic chips, enabling the scalable integration of computing resources to meet the growing demands of AI workloads. However, efficient inter-chiplet communication still faces significant bottlenecks, especially under the fine-grained and bursty Direct Memory Access (DMA) request patterns generated by processing elements in modern AI tasks. Existing communication models and simulators fail to capture these characteristics, which limits the accuracy of performance analysis and the effectiveness of optimization strategies. These limitations hinder DMA-communication inefficiencies in chiplet-based AI systems and pose challenges for designing HPC architectures. To address these challenges, we present the first comprehensive chiplet communication model that explicitly incorporates finegrained DMA traffic observed in realistic AI workloads. Building on this model, we propose COMET, a novel framework that intelligently searches for optimal DMA request aggregation and memory address mapping strategies tailored to chiplet environments. COMET dynamically consolidates small DMA transfers to improve bandwidth utilization and reduce communication latency, while also adapting on-chip memory mapping to align with workload-specific dataflows. This mitigates synchronization overhead across diverse AI tasks. Compared with inference on conventional chiplet communication schemes, COMET achieves1.1×−2.6×1.1 \times-2.6 \timesspeedup and1.5×−4.4×1.5 \times-4.4 \timeshigher bandwidth utilization across different DNN and LLM workloads.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖