FACE: Fully Overlapped PD Scheduling and Multi-Level Architecture Co-Exploration on Wafer
Zheng Xu, Dehao Kong, Jiaxin Liu, Dingcheng Jiang, Xu Dai, Jinyi Deng, Yang Hu, Shouyi Yin
摘要
The rapid expansion of large language models (LLMs) parameter scales imposes unprecedented demands on compute, memory, and communication resources for inference deployment. Wafer-scale chips, leveraging advanced packaging technologies, deliver high-density integration of compute and memory with high die-to-die (D2D) communication bandwidth, providing a compelling architectural approach to satisfy these resource requirements. However, its unprecedented chip area introduces significant architectural design complexities. Waferscale chips feature a multi-level architecture spanning the wafer, die, and core levels, involving numerous critical design parameters and trade-offs, which still lack systematic understanding and exploration. Moreover, this poses major challenges for LLM serving scheduling. Existing methods, largely adapted from GPUbased systems, fail to fully leverage the advantages of waferscale chips and mitigate their limitations, making it difficult to efficiently translate massive hardware resources into actual performance gains. To address these challenges, we introduce FACE, a coexploration framework for jointly optimizing multi-level architecture and serving scheduling. We first establish a flexible and extensible wafer-scale hardware template to systematically explore the optimal architecture and micro-architecture parameters. Leveraging the fine-grained control and high interconnect bandwidth of wafer-scale chips, FACE implements an LLM scheduling strategy that achieves fully overlapped prefill-decode execution and efficient KV cache management, maximizing hardware resource utilization to improve LLM service quality. Our evaluation demonstrates that FACE can achieve an average overall performance improvement of 3.68 × across various LLM models and datasets compared to the state-of-the-art (SOTA) LLM serving system on wafer-scale chips.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- WSC-LLM: Efficient LLM Service and Architecture Co-exploration for Wafer-scale ChipsZheng Xu, Dehao Kong, Jiaxin Liu, Jinxi Li 等ISCA 2025 · 被引用 22 次
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou 等HPCA 2026 · 被引用 2 次
- Cramming a Data Center into One Cabinet, a Co-Exploration of Computing and Hardware Architecture of Waferscale ChipXingmao Yu, Dingcheng Jiang, Jinyi Deng, Jingyao Liu 等ISCA 2025 · 被引用 11 次
- WaferLLM: Large Language Model Inference at Wafer ScaleCongjie He, Yeqi Huang, Pei Mu, Ziming Miao 等OSDI 2025 · 被引用 20 次
- Mapping and Communication Optimizations with Fault Tolerance for Wafer-Scale LLM InferenceJunwei Cui, Le Qin, Weilin Cai, Jiayi HuangISCA 2026
