Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event Notification
Tian Pan, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Ye Yang, Jiao Zhang
Abstract
Layer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation.
In this work, we present Hermes, a userspace-directed I/O event notification framework to enhance L7 LBs. Hermes uses userspace worker status to direct kernel-space connection dispatch. It implements lock-free concurrency management for inter-process worker status updates and retrievals, as well as scheduling decision synchronization from userspace to the kernel. In the kernel, Hermes leverages eBPF to non-intrusively override the reuseport socket selection for custom worker scheduling. Hermes has been deployed on O(100K) CPU cores in Alibaba Cloud, handling O(10M) RPS of traffic. It reduces daily worker hangs by 99.8% and lowers the unit cost of L7 LB infrastructure by 18.9%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05389f5d-3f81-4213-873e-425e623f8e41Cited by top-tier papers3
- Remote TCP Connection Offload and ApplicationsShuo Li, Steven W. D. Chien, Tianyi Gao, Michio HondaNSDI 2026 · 3 citations
- HybridMesh: A Hardware-software Hybrid Approach for Accelerating Service Mesh IngressMyoungsung You, Jaehyun Nam, Minjae Seo, Taejune Park et al.NSDI 2026 · 2 citations
- ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive ReroutingXiaoqing Sun, Xing Li, Xionglie Wei, Tian Pan et al.NSDI 2026
Builds on12
- AccelTCP: Accelerating Network Applications with Stateful TCP OffloadingYoungGyoun Moon, SeungEon Lee, Muhammad Asim Jamshed, KyoungSoo ParkNSDI 2020 · 121 citations
- Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switchesTian Pan, Nianbing Yu, Chenhao Jia, Jianwen Pi et al.SIGCOMM 2021 · 111 citations
- A High-Speed Load-Balancer Design with Guaranteed Per-Connection-ConsistencyTom Barbette, Chen Tang, Haoran Yao, Dejan Kostic et al.NSDI 2020 · 100 citations
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias et al.SOSP 2021 · 39 citations
- Canal Mesh: A Cloud-Scale Sidecar-Free Multi-Tenant Service Mesh ArchitectureEnge Song, Yang Song, Chengyun Lu, Tian Pan et al.SIGCOMM 2024 · 30 citations
Related papers
- Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed AsynchronyYifan Qiao, Chenxi Wang, Zhenyuan Ruan, Adam Belay et al.NSDI 2023 · 50 citations
- Capybara: Dynamic Load Balancing with Microsecond-Scale TCP MigrationInho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi et al.SIGCOMM 2026
- Scheduling Cloud Block Storage Proactively and Reactively with OmarXinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu et al.EuroSys 2026
- QDSR: Accelerating Layer-7 Load Balancing by Direct Server Return with QUICZiqi Wei, Zhiqiang Wang, Qing Li, Yuan Yang et al.USENIX ATC 2024 · 10 citations
- HA/TCP: A Reliable and Scalable Framework for TCP Network FunctionsHaoyu Gu, Ali José Mashtizadeh, Bernard WongNSDI 2025 · 1 citation
