Dynamic Load Balancer in Intel Xeon Scalable Processor: Performance Analyses, Enhancements, and Guidelines
Jiaqi Lou, Srikar Vanavasam, Yifan Yuan, Ren Wang, Nam Sung Kim
摘要
The rapid increase in inter-host networking speed has challenged host processing capabilities, as bursty traffic and uneven load distribution among host CPU cores give rise to excessive queuing delays and service latency variances. To cost-efficiently tackle this challenge, the latest Intel Xeon Scalable Processor has integrated an on-chip accelerator, named Dynamic Load Balancer (DLB). It consists of hardware queues, arbiters, and priority-based Quality of Service (QoS) features to maximize load-balancing performance while minimizing host CPU cycle consumption. In this work, we first compare the performance of DLB against popular softwarebased load balancers with a microbenchmark suite, demonstrating that DLB significantly outperforms them. Yet, DLB still consumes a significant number of host CPU cycles to prepare and enqueue work descriptors for received packets. Second, to eliminate the consumption of host CPU cycles in load balancing, we propose a system architecture/software co-design solution, AccDirect. Specifically, AccDirect leverages the Peer-to-Peer (P2P) communication capability of PCIe devices to enable a direct communication path between DLB and NIC, through which NIC directly enqueues the work descriptors. Our evaluation shows that AccDirect-based DLB offers practically the same performance as conventional DLB while reducing the system-wide power consumption of host CPU by 10%. Compared to the throughput of a commodity hardware-based load balancer for an end-to-end application, that of AccDirect-based DLB is 14-50% higher with comparable p99 latency. Lastly, we provide guidelines to make the best use of DLB, which is riddled with a vast configuration space and advanced features, after conducting a comprehensive evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SBB: Eliminating Centralized Bottlenecks in Userspace Network RuntimeKang Hu, Shuqi Dong, Chuandong Li, Ran Yi 等OSDI 2026
- ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling InsightsTao Lu, Jiapin Wang, Yelin Shan, Xiangping Zhang 等EuroSys 2026
它引用的顶会 Paper14
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper 等MICRO 2023 · 被引用 133 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 被引用 78 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
相关 Paper
- HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative ComputingJinghan Huang, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong 等ISCA 2024 · 被引用 7 次
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 被引用 88 次
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman 等ASPLOS 2026
- A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O DevicesHaneul Park, Jiaqi Lou, Sangjin Lee, Yifan Yuan 等ISCA 2025 · 被引用 2 次
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong 等MICRO 2025 · 被引用 3 次
