Reliable Actors with Retry Orchestration
Olivier Tardieu, David Grove, Gheorghe-Teodor Bercea, Paul Castro, Jaroslaw Cwiklik, Edward A. Epstein
摘要
Cloud developers have to build applications that are resilient to failures and interruptions. We advocate for a fault-tolerant programming model for the cloud based on actors, retry orchestration, and tail calls. This model builds upon persistent data stores and messages queues readily available on the cloud. Retry orchestration not only guarantees that (1) failed actor invocations will be retried but also that (2) completed invocations are never repeated and (3) it preserves a strict happen-before relationship across failures within call stacks. Tail calls can break complex tasks into simple steps to minimize re-execution during recovery. We review key application patterns and failure scenarios. We formalize a process calculus to precisely capture the mechanisms of fault tolerance in this model. We briefly describe our implementation. Using an application inspired by a typical enterprise scenario, we validate the functional correctness of our implementation and assess the impact of fault preparedness and recovery on performance.
CCS Concepts: • Software and its engineering → Error handling and recovery.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Durable functions: semantics for stateful serverlessSebastian Burckhardt, Chris Gillum, David Justo, Konstantinos Kallas 等OOPSLA 2021 · 被引用 63 次
- Fault-tolerant and transactional stateful serverless workflowsHaoran Zhang, Adney Cardoza, Peter Baile Chen, Sebastian Angel 等OSDI 2020 · 被引用 20 次
- Scalable and serializable networked multi-actor programmingBo Sang, Patrick Eugster, Gustavo Petri, Srivatsan Ravi 等OOPSLA 2020 · 被引用 3 次
相关 Paper
- Push-Button Reliability Testing for Cloud-Backed Applications with RainmakerYinfang Chen, Xudong Sun, Suman Nath, Ze Yang 等NSDI 2023 · 被引用 29 次
- Canary: Fault-Tolerant FaaS for Stateful Time-Sensitive ApplicationsMoiz Arif, Kevin Assogba, M. Mustafa RafiqueSC 2022 · 被引用 9 次
- A fault-tolerance shim for serverless computingVikram Sreekanti, Chenggang Wu, Saurav Chhatrapati, Joseph E. Gonzalez 等EuroSys 2020 · 被引用 58 次
- If At First You Don't Succeed, Try, Try, Again...? Insights and LLM-informed Tooling for Detecting Retry Bugs in Software SystemsBogdan Alexandru Stoica, Utsav Sethi, Yiming Su, Cyrus Zhou 等SOSP 2024 · 被引用 4 次
- Coverage Guided Fault Injection for Cloud SystemsYu Gao, Wensheng Dou, Dong Wang, Wenhan Feng 等ICSE 2023 · 被引用 13 次
