MeshRT: Compile-Time Governed Wafer-Scale Runtime for Low-Latency High-Throughput Inference
Congjie He, Le Xu, Zhan Lu, Yeqi Huang, Haocheng Xiao, Cheng Deng, Lingxiao Ma, Ziming Miao, Fan Yang, Luo Mai
Abstract
Wafer-scale accelerators promise ultra-low-latency AI inference, but current system stacks still carry an unsustainably high cost premium. The reason is that many current and emerging inference techniques, such as batching and MoE, introduce runtime dynamism that existing wafer-scale systems cannot support efficiently. As a result, they must choose between two unsatisfactory options: accept suboptimal latency, or preserve low latency by restricting dynamic techniques, thereby limiting the models they can support and reducing aggregate throughput. This tradeoff significantly increases the cost of low-latency serving.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ce315b56-eefd-4120-9a7f-d39cca3f9ab2Related papers
- Wavel: A Fast and Efficient Compilation System for Wafer-Scale AcceleratorsYeqi Huang, Congjie He, Haocheng Xiao, Yanwei Ye et al.SOSP 2026
- High-throughput and Flexible Host Networking for Accelerated ComputingAthinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai et al.OSDI 2024 · 11 citations
- FACE: Fully Overlapped PD Scheduling and Multi-Level Architecture Co-Exploration on WaferZheng Xu, Dehao Kong, Jiaxin Liu, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismRuidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu et al.SIGCOMM 2025 · 19 citations
- WaferLLM: Large Language Model Inference at Wafer ScaleCongjie He, Yeqi Huang, Pei Mu, Ziming Miao et al.OSDI 2025 · 20 citations
