Lune

HPCA2026Top-tier venue

FACE: Fully Overlapped PD Scheduling and Multi-Level Architecture Co-Exploration on Wafer

Zheng Xu, Dehao Kong, Jiaxin Liu, Dingcheng Jiang, Xu Dai, Jinyi Deng, Yang Hu, Shouyi Yin

2026Year
2Citations

Abstract

The rapid expansion of large language models (LLMs) parameter scales imposes unprecedented demands on compute, memory, and communication resources for inference deployment. Wafer-scale chips, leveraging advanced packaging technologies, deliver high-density integration of compute and memory with high die-to-die (D2D) communication bandwidth, providing a compelling architectural approach to satisfy these resource requirements. However, its unprecedented chip area introduces significant architectural design complexities. Waferscale chips feature a multi-level architecture spanning the wafer, die, and core levels, involving numerous critical design parameters and trade-offs, which still lack systematic understanding and exploration. Moreover, this poses major challenges for LLM serving scheduling. Existing methods, largely adapted from GPUbased systems, fail to fully leverage the advantages of waferscale chips and mitigate their limitations, making it difficult to efficiently translate massive hardware resources into actual performance gains. To address these challenges, we introduce FACE, a coexploration framework for jointly optimizing multi-level architecture and serving scheduling. We first establish a flexible and extensible wafer-scale hardware template to systematically explore the optimal architecture and micro-architecture parameters. Leveraging the fine-grained control and high interconnect bandwidth of wafer-scale chips, FACE implements an LLM scheduling strategy that achieves fully overlapped prefill-decode execution and efficient KV cache management, maximizing hardware resource utilization to improve LLM service quality. Our evaluation demonstrates that FACE can achieve an average overall performance improvement of 3.68 × across various LLM models and datasets compared to the state-of-the-art (SOTA) LLM serving system on wafer-scale chips.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get d25fc603-ea0c-43f0-bd5a-8093c8a1e5bd

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines