Fault Tolerant Service Function Chaining
Milad Ghaznavi, Elaheh Jalalpour, Bernard Wong, Raouf Boutaba, Ali José Mashtizadeh
Abstract
Network traffic typically traverses a sequence of middleboxes forming a service function chain, or simply a chain. Tolerating failures when they occur along chains is imperative to the availability and reliability of enterprise applications. Making a chain fault-tolerant is challenging since, in the event of failures, the state of faulty middleboxes must be correctly and quickly recovered while providing high throughput and low latency.
In this paper, we introduce FTC, a system design and protocol for fault-tolerant service function chaining. FTC provides strong consistency with up to f middlebox failures for chains of length f + 1 or longer without requiring dedicated replica nodes. In FTC, state updates caused by packet processing at a middlebox are collected, piggybacked onto the packet, and sent along the chain to be replicated. Our evaluation shows that compared with the state of art [51], FTC improves throughput by 2-3.5× for a chain of two to five middleboxes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 131ba13a-a639-4c03-8ef0-e415a806bdaaCited by top-tier papers2
- A composition framework for change managementAjay Mahimkar, Carlos Eduardo de Andrade, Rakesh K. Sinha, Giritharan RanaSIGCOMM 2021 · 20 citations
- HA/TCP: A Reliable and Scalable Framework for TCP Network FunctionsHaoyu Gu, Ali José Mashtizadeh, Bernard WongNSDI 2025 · 1 citation
Related papers
- High Throughput Replication with Integrated Membership ManagementPedro Fouto, Nuno M. Preguiça, João LeitãoUSENIX ATC 2022
- Meerkat: Scalable, Network-Aware Failure Recovery for the Internet of ThingsAnastasiia Kozar, Ankit Chaudhary, Steffen Zeuch, Volker MarklVLDB 2026
- Ambulance: Saving BFT through RacingNeil Giridharan, Shubham Mishra, Lorenzo Alvisi, Natacha Crooks et al.OSDI 2026 · 1 citation
- Avicenna: Masking Slowdowns in Replicated State Machines with Counterfactual EvaluationChristopher Hodsdon, Zijian Qin, Khiem Ngo, Siddhartha Sen et al.EuroSys 2026
- HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter servicesMarios Kogias, Edouard BugnionEuroSys 2020 · 52 citations
