Re-architecting End-host Networking with CXL: Coherence, Memory, and Offloading
Houxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong, Ren Wang, Saksham Agarwal, Nam Sung Kim
Abstract
The traditional Network Interface Controller (NIC) suffers from the inherent inefficiency of the PCIe interconnect with two key limitations. First, since it allows the NIC to transfer packets to the host CPU memory only through DMA, it incurs high latency, the impact of which becomes more pronounced, especially for smallsized packets. Second, it supports neither shared memory nor full cache coherence between the host CPU and the NIC. Therefore, the host CPU can access NIC memory only through MMIO-which provides higher latency and lower bandwidth than cache-coherent memory access-and software is often responsible for managing consistency between the host CPU and NIC memory. Although built on the PCIe interconnect, Compute Express Link (CXL) efficiently addresses these limitations by providing hardware-managed unified memory and cache coherence between the host CPU and NIC memory. This allows the host CPU and the NIC to access each other's memory using load/store semantics, offering low latency and high bandwidth. In this work, we first present a Type-1 CXL-NIC design that replaces slow legacy PCIe transactions with fast CXL.cache transactions for NIC-to-CPU memory accesses. Second, we extend the Type-1 CXL-NIC to a Type-2 CXL-NIC that introduces cachecoherent NIC memory exposed to the host CPU through CXL.mem, which can buffer packets and descriptors 1 . Lastly, we demonstrate a networking-application co-acceleration by exploiting unique CXL 1 Descriptors store information needed for packet processing, including packet sizes and pointers to DMA ring buffers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55dde1c4-5240-4dde-989e-19185afa93f6Cited by top-tier papers1
Ask how each one uses itBuilds on25
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- A Software Approach to Defeating Side Channels in Last-Level CachesZiqiao Zhou, Michael K. Reiter, Yinqian ZhangCCS 2016 · 155 citations
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper et al.MICRO 2023 · 133 citations
- AccelTCP: Accelerating Network Applications with Stateful TCP OffloadingYoungGyoun Moon, SeungEon Lee, Muhammad Asim Jamshed, KyoungSoo ParkNSDI 2020 · 121 citations
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im et al.SOSP 2021 · 83 citations
Related papers
- CC-NIC: a Cache-Coherent Interface to the NICHenry N. Schuh, Arvind Krishnamurthy, David E. Culler, Henry M. Levy et al.ASPLOS 2024 · 19 citations
- CTXNL: A Software-Hardware Co-designed Solution for Efficient CXL-Based Transaction ProcessingZhao Wang, Yiqi Chen, Cong Li, Yijin Guan et al.ASPLOS 2025 · 9 citations
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin et al.MICRO 2024 · 26 citations
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman et al.ASPLOS 2026
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang et al.HPCA 2026 · 1 citation
