MeshRT: Compile-Time Governed Wafer-Scale Runtime for Low-Latency High-Throughput Inference
Congjie He, Le Xu, Zhan Lu, Yeqi Huang, Haocheng Xiao, Cheng Deng, Lingxiao Ma, Ziming Miao, Fan Yang, Luo Mai
摘要
Wafer-scale accelerators promise ultra-low-latency AI inference, but current system stacks still carry an unsustainably high cost premium. The reason is that many current and emerging inference techniques, such as batching and MoE, introduce runtime dynamism that existing wafer-scale systems cannot support efficiently. As a result, they must choose between two unsatisfactory options: accept suboptimal latency, or preserve low latency by restricting dynamic techniques, thereby limiting the models they can support and reducing aggregate throughput. This tradeoff significantly increases the cost of low-latency serving.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Wavel: A Fast and Efficient Compilation System for Wafer-Scale AcceleratorsYeqi Huang, Congjie He, Haocheng Xiao, Yanwei Ye 等SOSP 2026
- High-throughput and Flexible Host Networking for Accelerated ComputingAthinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai 等OSDI 2024 · 被引用 11 次
- FACE: Fully Overlapped PD Scheduling and Multi-Level Architecture Co-Exploration on WaferZheng Xu, Dehao Kong, Jiaxin Liu, Dingcheng Jiang 等HPCA 2026 · 被引用 2 次
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismRuidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu 等SIGCOMM 2025 · 被引用 19 次
- WaferLLM: Large Language Model Inference at Wafer ScaleCongjie He, Yeqi Huang, Pei Mu, Ziming Miao 等OSDI 2025 · 被引用 20 次
