Rose: Reproducing External-Fault-Induced Failures in Distributed Systems with Lightweight Instrumentation
Sebastião Amaro, Pedro Fonseca, Miguel Matos
摘要
Distributed systems form the backbone of critical infrastructures, yet remain vulnerable to external-fault-induced bugs that manifest only when specific external events occur during specific application states. Existing approaches to reproducing these bugs require fine-grained information about the application, which might not be available in production, and operate within a limited fault model. Rose is a novel approach that collects traces from production and systematically generates fault schedules that reproduce these bugs. By leveraging the insight that external faults are observable through system interfaces, Rose uses lightweight tracing (2.6% overhead) to capture essential application-environment interactions. Then, it identifies the application states when faults must occur to trigger bugs, and generates schedules that consistently reproduce these bugs. Rose reproduced 20 bugs across eight production systems implemented in diverse languages (C, C++, Java, Go, Scala), including widely-used systems such as Zookeeper, MongoDB, and HBase.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Coverage Guided Fault Injection for Cloud SystemsYu Gao, Wensheng Dou, Dong Wang, Wenhan Feng 等ICSE 2023 · 被引用 13 次
- Efficient Reproduction of Fault-Induced Failures in Distributed Systems with Feedback-Driven Fault InjectionJia Pan, Haoze Wu, Tanakorn Leesatapornwongsa, Suman Nath 等SOSP 2024 · 被引用 4 次
- Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract StatesHaoze Wu, Jia Pan, Peng HuangNSDI 2024 · 被引用 15 次
- When Amnesia Strikes: Understanding and Reproducing Data Loss Bugs with Fault InjectionMaria Ramos, João Azevedo, Kyle Kingsbury, José Pereira 等VLDB 2024 · 被引用 2 次
- TailTracer: Continuous Tail Tracing for Production UseTianyi Liu, Yi Li, Yiyu Zhang, Zhuangda Wang 等OOPSLA 2025 · 被引用 1 次
