Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load Balancing
Chaoliang Zeng, Layong Luo, Teng Zhang, Zilong Wang, Luyang Li, Wenchen Han, Nan Chen, Lebing Wan, Lichao Liu, Zhipeng Ding, Xiongfei Geng, Tao Feng
Abstract
Stateful layer-4 load balancers (LB) are deployed at datacenter boundaries to distribute Internet traffic to backend real servers. To steer terabits per second traffic, traditional software LBs scale out with many expensive servers. Recent switch-accelerated LBs scale up efficiently, but fail to offload a massive number of concurrent flows into limited on-chip SRAMs.
This paper presents Tiara, a hardware architecture for stateful layer-4 LBs that aims to support a high traffic rate (> 1 Tbps), a large number of concurrent flows (> 10M), and many new connections per second (> 1M), without any assumption on traffic patterns. The three-tier architecture of Tiara makes the best use of heterogeneous hardware for stateful LBs, including a programmable switch and FPGAs for the fast path and x86 servers for the slow path. The core idea of Tiara is to divide the LB fast path into a memory-intensive task (real server selection) and a throughput-intensive task (packet encap/decap), and map them into the most suitable hardware, respectively (i.e., map real server selection into FPGA with large high-bandwidth memory (HBM) and packet encap/decap into a high-throughput programmable switch). We have implemented a fully functional Tiara prototype, and experiments show that Tiara can achieve extremely high performance (1.6 Tbps throughput, 80M concurrent flows, 1.8M new connections per second, and less than 4 us latency in the fast path) in a holistic server equipped with 8 FPGA cards, with high cost, energy, and space efficiency.
- This work is done while Chaoliang Zeng, Zilong Wang, Luyang Li, and Wenchen Han are interns in Bytedance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aaadf041-bfc8-44de-8f7f-d5622f67ad9cCited by top-tier papers24
- ClickINC: In-network Computing as a Service in Heterogeneous Programmable Data-center NetworksWenquan Xu, Zijian Zhang, Yong Feng, Haoyu Song et al.SIGCOMM 2023 · 34 citations
- Disaggregating Stateful Network FunctionsDeepak Bansal, Gerald DeGrace, Rishabh Tewari, Michal Zygmunt et al.NSDI 2023 · 33 citations
- FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated LearningJunxue Zhang, Xiaodian Cheng, Wei Wang, Liu Yang et al.NSDI 2023 · 29 citations
- Unleashing SmartNIC Packet Processing Performance in P4Jiarong Xing, Yiming Qiu, Kuo-Feng Hsu, Songyuan Sui et al.SIGCOMM 2023 · 29 citations
- LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge CloudsTian Pan, Kun Liu, Xionglie Wei, Yisong Qiao et al.NSDI 2024 · 27 citations
Builds on3
- TEA: Enabling State-Intensive Network Functions on Programmable SwitchesDaehyeok Kim, Zaoxing Liu, Yibo Zhu, Changhoon Kim et al.SIGCOMM 2020 · 121 citations
- A High-Speed Load-Balancer Design with Guaranteed Per-Connection-ConsistencyTom Barbette, Chen Tang, Haoran Yao, Dejan Kostic et al.NSDI 2020 · 100 citations
- FPGA-Accelerated Compactions for LSM-based Key-Value StoreTeng Zhang, Jianying Wang, Xuntao Cheng, Hao Xu et al.FAST 2020 · 99 citations
Related papers
- Capybara: Dynamic Load Balancing with Microsecond-Scale TCP MigrationInho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi et al.SIGCOMM 2026
- Miresga: Accelerating Layer-7 Load Balancing with Programmable SwitchesXiaoyi Shi, Lin He, Jiasheng Zhou, Yifan Yang et al.WWW 2025 · 3 citations
- Scalable On-Switch Rate Limiters for the CloudYongchao He, Wenfei Wu, Xuemin Wen, Haifeng Li et al.INFOCOM 2021 · 14 citations
- FAERY: An FPGA-accelerated Embedding-based Retrieval SystemChaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han et al.OSDI 2022 · 9 citations
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng et al.NSDI 2023 · 154 citations
