CC-NIC: a Cache-Coherent Interface to the NIC
Henry N. Schuh, Arvind Krishnamurthy, David E. Culler, Henry M. Levy, Luigi Rizzo, Samira Manabi Khan, Brent E. Stephens
摘要
Emerging interconnects make peripherals, such as the network interface controller (NIC), accessible through the processor's cache hierarchy, allowing these devices to participate in the CPU cache coherence protocol. This is a fundamental change from the separate I/O data paths and read-write transaction primitives of today's PCIe NICs. Our experiments show that the I/O data path characteristics cause NICs to prioritize CPU efficiency at the expense of inflated latency, an issue that can be mitigated by the emerging low-latency coherent interconnects. But, the coherence abstraction is not suited to current host-NIC access patterns. Applying existing signaling mechanisms and data structure layouts in a cache-coherent setting results in extraneous communication and cache retention, limiting performance. Redesigning the interface is necessary to minimize overheads and benefit from the new interactions coherence enables. This work contributes CC-NIC, a host-NIC interface design for coherent interconnects. We model CC-NIC using Intel's Ice Lake and Sapphire Rapids UPI interconnects, demonstrating the potential of optimizing for coherence. Our results show a maximum packet rate of 1.5Gpps and 980Gbps packet throughput. CC-NIC has 77% lower minimum latency, and 88% lower at 80% load, than today's PCIe NICs. We also demonstrate application-level core savings. Finally, we show that CC-NIC's benefits hold across a range of interconnect performance characteristics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Accelerating Retrieval-Augmented GenerationDerrick Quinn, Mohammad Nouri, Neel Patel, John Salihu 等ASPLOS 2025 · 被引用 37 次
- CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsBowen Liu, Xinyang Huang, Qijing Li, Zhuobin Huang 等SIGCOMM 2025 · 被引用 10 次
- HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative ComputingJinghan Huang, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong 等ISCA 2024 · 被引用 7 次
- DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack CommunicationXu Zhang, Ke Liu, Yuan Hui, Xiaolong Zheng 等USENIX ATC 2025 · 被引用 5 次
- Opening Up Kernel-Bypass TCP StacksShinichi Awamoto, Michio HondaUSENIX ATC 2025 · 被引用 5 次
它引用的顶会 Paper9
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper 等MICRO 2023 · 被引用 133 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- Xenic: SmartNIC-Accelerated Distributed TransactionsHenry N. Schuh, Weihao Liang, Ming Liu, Jacob Nelson 等SOSP 2021 · 被引用 62 次
- Dagger: efficient and fast RPCs in cloud microservices with near-memory reconfigurable NICsNikita Lazarev, Shaojie Xiang, Neil Adit, Zhiru Zhang 等ASPLOS 2021 · 被引用 54 次
- Fine-Grained Isolation for Scalable, Dynamic, Multi-tenant Edge CloudsYuxin Ren, Guyue Liu, Vlad Nitu, Wenyuan Shao 等USENIX ATC 2020 · 被引用 47 次
相关 Paper
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong 等MICRO 2025 · 被引用 3 次
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman 等ASPLOS 2026
- Rambda: RDMA-driven Acceleration Framework for Memory-intensive µs-scale Datacenter ApplicationsYifan Yuan, Jinghan Huang, Yan Sun, Tianchen Wang 等HPCA 2023 · 被引用 25 次
- Cross-Core Interrupt Detection: Exploiting User and Virtualized IPIsFabian Rauscher, Daniel GrussCCS 2024 · 被引用 4 次
- Dynamic Load Balancer in Intel Xeon Scalable Processor: Performance Analyses, Enhancements, and GuidelinesJiaqi Lou, Srikar Vanavasam, Yifan Yuan, Ren Wang 等ISCA 2025 · 被引用 4 次
