R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic System
Kefei Liu, Zhuo Jiang, Jiao Zhang, Shixian Guo, Xuan Zhang, Yangyang Bai, Yongbin Dong, Feng Luo, Zhang Zhang, Lei Wang, Xiang Shi, Haohan Xu
摘要
RoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related.
In this paper, we propose R-Pingmesh, the first service-aware RoCE network monitoring and diagnostic system based on endto-end probing. R-Pingmesh can accurately measure network RTT and end-host processing delay based on commodity RDMA NICs (RNICs), distinguish between RNIC and in-network packet drops, and judge whether a problem is network-related and assess its impact on services. We have deployed R-Pingmesh on tens of thousands of RNICs for over 6 months. One-month evaluation results show that 85% of the problems located by R-Pingmesh are accurate, where all 157 switch network problems are accurate. R-Pingmesh efficiently detects and locates 14 types of problems during deployment, and we share our experience in dealing with them.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang 等NSDI 2025 · 被引用 36 次
- Evolution of Aegis: Fault Diagnosis for AI Model Training Service in ProductionJianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng 等NSDI 2025 · 被引用 21 次
- Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleQingkai Meng, Hao Zheng, Zhenhui Zhang, ChonLam Lao 等SIGCOMM 2025 · 被引用 16 次
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin 等SIGCOMM 2025 · 被引用 13 次
- SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model TrainingWei Liu, Kun Qian, Zhenhua Li, Tianyin Xu 等SIGCOMM 2025 · 被引用 8 次
它引用的顶会 Paper15
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 等NSDI 2024 · 被引用 415 次
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- PINT: Probabilistic In-band Network TelemetryRan Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi 等SIGCOMM 2020 · 被引用 268 次
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi 等NSDI 2021 · 被引用 228 次
- Flow Event Telemetry on Programmable Data PlaneYu Zhou, Chen Sun, Hongqiang Harry Liu, Rui Miao 等SIGCOMM 2020 · 被引用 139 次
相关 Paper
- Hostping: Diagnosing Intra-host Network Bottlenecks in RDMA ServersKefei Liu, Zhuo Jiang, Jiao Zhang, Haoran Wei 等NSDI 2023 · 被引用 52 次
- ByteTracker: An Agentless and Real-time Path-aware Network Probing SystemShixian Guo, Kefei Liu, Yulin Lai, Yangyang Bai 等SIGCOMM 2025 · 被引用 8 次
- INSERT: In-Network Stateful End-to-End RDMA TelemetryHyunseok Chang, Walid A. Hanafy, Sarit Mukherjee, Limin WangINFOCOM 2024 · 被引用 3 次
- Mitigating Scalability Walls of RDMA-based Container NetworksWei Liu, Kun Qian, Zhenhua Li, Feng Qian 等NSDI 2025 · 被引用 9 次
- Hawkeye: Diagnosing RDMA Network Performance Anomalies with PFC ProvenanceShicheng Wang, Menghao Zhang, Xiao Li, Qiyang Peng 等SIGCOMM 2025 · 被引用 7 次
