Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load Balancing
Chaoliang Zeng, Layong Luo, Teng Zhang, Zilong Wang, Luyang Li, Wenchen Han, Nan Chen, Lebing Wan, Lichao Liu, Zhipeng Ding, Xiongfei Geng, Tao Feng
摘要
Stateful layer-4 load balancers (LB) are deployed at datacenter boundaries to distribute Internet traffic to backend real servers. To steer terabits per second traffic, traditional software LBs scale out with many expensive servers. Recent switch-accelerated LBs scale up efficiently, but fail to offload a massive number of concurrent flows into limited on-chip SRAMs.
This paper presents Tiara, a hardware architecture for stateful layer-4 LBs that aims to support a high traffic rate (> 1 Tbps), a large number of concurrent flows (> 10M), and many new connections per second (> 1M), without any assumption on traffic patterns. The three-tier architecture of Tiara makes the best use of heterogeneous hardware for stateful LBs, including a programmable switch and FPGAs for the fast path and x86 servers for the slow path. The core idea of Tiara is to divide the LB fast path into a memory-intensive task (real server selection) and a throughput-intensive task (packet encap/decap), and map them into the most suitable hardware, respectively (i.e., map real server selection into FPGA with large high-bandwidth memory (HBM) and packet encap/decap into a high-throughput programmable switch). We have implemented a fully functional Tiara prototype, and experiments show that Tiara can achieve extremely high performance (1.6 Tbps throughput, 80M concurrent flows, 1.8M new connections per second, and less than 4 us latency in the fast path) in a holistic server equipped with 8 FPGA cards, with high cost, energy, and space efficiency.
- This work is done while Chaoliang Zeng, Zilong Wang, Luyang Li, and Wenchen Han are interns in Bytedance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- ClickINC: In-network Computing as a Service in Heterogeneous Programmable Data-center NetworksWenquan Xu, Zijian Zhang, Yong Feng, Haoyu Song 等SIGCOMM 2023 · 被引用 34 次
- Disaggregating Stateful Network FunctionsDeepak Bansal, Gerald DeGrace, Rishabh Tewari, Michal Zygmunt 等NSDI 2023 · 被引用 33 次
- FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated LearningJunxue Zhang, Xiaodian Cheng, Wei Wang, Liu Yang 等NSDI 2023 · 被引用 29 次
- Unleashing SmartNIC Packet Processing Performance in P4Jiarong Xing, Yiming Qiu, Kuo-Feng Hsu, Songyuan Sui 等SIGCOMM 2023 · 被引用 29 次
- LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge CloudsTian Pan, Kun Liu, Xionglie Wei, Yisong Qiao 等NSDI 2024 · 被引用 27 次
它引用的顶会 Paper3
- TEA: Enabling State-Intensive Network Functions on Programmable SwitchesDaehyeok Kim, Zaoxing Liu, Yibo Zhu, Changhoon Kim 等SIGCOMM 2020 · 被引用 121 次
- A High-Speed Load-Balancer Design with Guaranteed Per-Connection-ConsistencyTom Barbette, Chen Tang, Haoran Yao, Dejan Kostic 等NSDI 2020 · 被引用 100 次
- FPGA-Accelerated Compactions for LSM-based Key-Value StoreTeng Zhang, Jianying Wang, Xuntao Cheng, Hao Xu 等FAST 2020 · 被引用 99 次
相关 Paper
- Capybara: Dynamic Load Balancing with Microsecond-Scale TCP MigrationInho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi 等SIGCOMM 2026
- Miresga: Accelerating Layer-7 Load Balancing with Programmable SwitchesXiaoyi Shi, Lin He, Jiasheng Zhou, Yifan Yang 等WWW 2025 · 被引用 3 次
- Scalable On-Switch Rate Limiters for the CloudYongchao He, Wenfei Wu, Xuemin Wen, Haifeng Li 等INFOCOM 2021 · 被引用 14 次
- FAERY: An FPGA-accelerated Embedding-based Retrieval SystemChaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han 等OSDI 2022 · 被引用 9 次
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng 等NSDI 2023 · 被引用 154 次
