Demystifying a CXL Type-2 Device: A Heterogeneous Cooperative Computing Perspective
Houxiang Ji, Srikar Vanavasam, Yang Zhou, Qirong Xia, Jinghan Huang, Yifan Yuan, Ren Wang, Pekon Gupta, Bhushan Chitlur, Ipoom Jeong, Nam Sung Kim
Abstract
CXL is the latest interconnect technology built on PCIe, providing three protocols to facilitate three distinct types of devices, each with unique capabilities. Among these devices, a CXL Type-2 device has become commercially available, followed by CXL Type-3 devices. Therefore, it is timely to understand capabilities and characteristics of the CXL Type-2 device, as well as explore suitable applications. In this work, first, we delve into three key features of a CXL Type-2 device: cache-coherent device accelerator to host memory, device accelerator to device memory, and host CPU to device memory accesses. Second, using microbenchmarks, we comprehensively characterize the latency and bandwidth of these memory accesses with a CXL Type-2 device, and then compare them with those of equivalent memory accesses with comparable devices, such as emulated CXL Type-2, CXL Type-3, and PCIe devices. Lastly, as applications that exploit the unique capabilities of a CXL Type-2 device, we propose two CXL-based Linux memory optimization features: compressed RAM cache for swap (zswap) and memory deduplication (ksm). Our evaluation shows that Redis, when running with traditional CPU-based zswap and ksm, suffers from a tail latency increase of 4.5-10.3× compared to Redis running alone. While PCIe-based zswap and ksm still experience a tail latency increase of up to 8.1×, CXL-based zswap and ksm practically eliminate the tail latency increase with faster and more efficient host-device communication than PCIe-based zswap and ksm.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 13f60d08-2432-441d-80df-61c03ff353fcCited by top-tier papers11
- Tigon: A Distributed Database for a CXL PodYibo Huang, Haowei Chen, Newton Ni, Yan Sun et al.OSDI 2025 · 12 citations
- Para-ksm: Parallelized Memory Deduplication with Data Streaming AcceleratorHouxiang Ji, Minho Kim, Seonmu Oh, Daehoon Kim et al.USENIX ATC 2025 · 7 citations
- Dynamic Load Balancer in Intel Xeon Scalable Processor: Performance Analyses, Enhancements, and GuidelinesJiaqi Lou, Srikar Vanavasam, Yifan Yuan, Ren Wang et al.ISCA 2025 · 4 citations
- A Programming Model for Disaggregated Memory over CXLGal Assa, Moritz Lumme, Lucas Bürgi, Michal Friedman et al.ASPLOS 2026 · 4 citations
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong et al.MICRO 2025 · 3 citations
Related papers
- STYX: Exploiting SmartNIC Capability to Reduce Datacenter Memory TaxHouxiang Ji, Mark Mansi, Yan Sun, Yifan Yuan et al.USENIX ATC 2023 · 22 citations
- AXLE: Coordinated Offloading with Asynchronous Back-Streaming in Computational Memory SystemsSuyeon Lee, Kangkyu Park, Kwangsik Shin, Ada GavrilovskaISCA 2026
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper et al.MICRO 2023 · 133 citations
- Cxlalloc: Safe and Efficient Memory Allocation for a CXL PodNewton Ni, Yan Sun, Zhiting Zhu, Emmett WitchelASPLOS 2026 · 2 citations
- NeoMem: Hardware/Software Co-Design for CXL-Native Memory TieringZhe Zhou, Yiqi Chen, Tao Zhang, Yang Wang et al.MICRO 2024 · 17 citations
