R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic System
Kefei Liu, Zhuo Jiang, Jiao Zhang, Shixian Guo, Xuan Zhang, Yangyang Bai, Yongbin Dong, Feng Luo, Zhang Zhang, Lei Wang, Xiang Shi, Haohan Xu
Abstract
RoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related.
In this paper, we propose R-Pingmesh, the first service-aware RoCE network monitoring and diagnostic system based on endto-end probing. R-Pingmesh can accurately measure network RTT and end-host processing delay based on commodity RDMA NICs (RNICs), distinguish between RNIC and in-network packet drops, and judge whether a problem is network-related and assess its impact on services. We have deployed R-Pingmesh on tens of thousands of RNICs for over 6 months. One-month evaluation results show that 85% of the problems located by R-Pingmesh are accurate, where all 157 switch network problems are accurate. R-Pingmesh efficiently detects and locates 14 types of problems during deployment, and we share our experience in dealing with them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30637674-4da0-4c13-9736-ea5cc55d490cCited by top-tier papers12
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang et al.NSDI 2025 · 36 citations
- Evolution of Aegis: Fault Diagnosis for AI Model Training Service in ProductionJianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng et al.NSDI 2025 · 21 citations
- Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleQingkai Meng, Hao Zheng, Zhenhui Zhang, ChonLam Lao et al.SIGCOMM 2025 · 16 citations
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin et al.SIGCOMM 2025 · 13 citations
- SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model TrainingWei Liu, Kun Qian, Zhenhua Li, Tianyin Xu et al.SIGCOMM 2025 · 8 citations
Builds on15
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- PINT: Probabilistic In-band Network TelemetryRan Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi et al.SIGCOMM 2020 · 268 citations
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi et al.NSDI 2021 · 228 citations
- Flow Event Telemetry on Programmable Data PlaneYu Zhou, Chen Sun, Hongqiang Harry Liu, Rui Miao et al.SIGCOMM 2020 · 139 citations
Related papers
- Hostping: Diagnosing Intra-host Network Bottlenecks in RDMA ServersKefei Liu, Zhuo Jiang, Jiao Zhang, Haoran Wei et al.NSDI 2023 · 52 citations
- ByteTracker: An Agentless and Real-time Path-aware Network Probing SystemShixian Guo, Kefei Liu, Yulin Lai, Yangyang Bai et al.SIGCOMM 2025 · 8 citations
- INSERT: In-Network Stateful End-to-End RDMA TelemetryHyunseok Chang, Walid A. Hanafy, Sarit Mukherjee, Limin WangINFOCOM 2024 · 3 citations
- Mitigating Scalability Walls of RDMA-based Container NetworksWei Liu, Kun Qian, Zhenhua Li, Feng Qian et al.NSDI 2025 · 9 citations
- Hawkeye: Diagnosing RDMA Network Performance Anomalies with PFC ProvenanceShicheng Wang, Menghao Zhang, Xiao Li, Qiyang Peng et al.SIGCOMM 2025 · 7 citations
