Collie: Finding Performance Anomalies in RDMA Subsystems
Xinhao Kong, Yibo Zhu, Huaping Zhou, Zhuo Jiang, Jianxi Ye, Chuanxiong Guo, Danyang Zhuo
摘要
High-speed RDMA networks are getting rapidly adopted in the industry for their low latency and reduced CPU overheads. To verify that RDMA can be used in production, system administrators need to understand the set of application workloads that can potentially trigger abnormal performance behaviors (e.g., unexpected low throughput, PFC pause frame storm). We design and implement Collie, a tool for users to systematically uncover performance anomalies in RDMA subsystems without the need to access hardware internal designs. Instead of individually testing each hardware device (e.g., NIC, memory, PCIe), Collie is holistic, constructing a comprehensive search space for application workloads. Collie then uses simulated annealing to drive RDMA-related performance and diagnostic counters to extreme value regions to find workloads that can trigger performance anomalies. We evaluate Collie on combinations of various RDMA NIC, CPU, and other hardware components. Collie found 15 new performance anomalies. All of them are acknowledged by the hardware vendors. 7 of them are already fixed after we reported them. We also present our experience in using Collie to avoid performance anomalies for an RDMA RPC library and an RDMA distributed machine learning framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang 等NSDI 2024 · 被引用 192 次
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng 等NSDI 2023 · 被引用 154 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu 等NSDI 2025 · 被引用 82 次
- Understanding RDMA Microarchitecture Resources for Performance IsolationXinhao Kong, Jingrong Chen, Wei Bai, Yechen Xu 等NSDI 2023 · 被引用 81 次
它引用的顶会 Paper7
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- kAFL: Hardware-Assisted Feedback Fuzzing for OS KernelsSergej Schumilo, Cornelius Aschermann, Robert Gawlik, Sebastian Schinzel 等USENIX Security 2017 · 被引用 324 次
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi 等NSDI 2021 · 被引用 228 次
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 被引用 88 次
- RDMA is Turing complete, we just did not know it yet!Waleed Reda, Marco Canini, Dejan Kostic, Simon PeterNSDI 2022 · 被引用 56 次
相关 Paper
- Vedrfolnir: RDMA Network Performance Anomalies Diagnosis in Collective CommunicationsYuxuan Chen, Menghao Zhang, Xiheng Li, Fangzheng Jiao 等INFOCOM 2026 · 被引用 1 次
- Hawkeye: Diagnosing RDMA Network Performance Anomalies with PFC ProvenanceShicheng Wang, Menghao Zhang, Xiao Li, Qiyang Peng 等SIGCOMM 2025 · 被引用 7 次
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun 等PPoPP 2026
- Maximizing the Benefit of RDMA at End HostsXiaoliang Wang, Hexiang Song, Cam-Tu Nguyen, Dongxu Cheng 等INFOCOM 2021 · 被引用 8 次
- UCCL-Tran: An Extensible Software Transport Layer for GPU NetworkingYang Zhou, Zhongjie Chen, Ziming Mao, ChonLam Lao 等OSDI 2026
