Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders
Hyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Euicheol Lim, Gwangsun Kim
Abstract
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not costeffective for NDP because they are not optimized for memorybound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur µs-scale latency and are not suitable for fine-grained NDP.
To achieve high-performance NDP end-to-end, we propose a low-overhead general-purpose NDP architecture for CXL memory referred to as Memory-Mapped NDP (M 2 NDP), which comprises memory-mapped functions (M 2 func) and memory-mapped µthreading (M 2 µthread). M 2 func is a CXL.mem-compatible lowoverhead communication mechanism between the host processor and NDP controller in CXL memory. M 2 µthread enables lowcost, general-purpose NDP unit design by introducing lightweight µthreads that support highly concurrent execution of kernels with minimal resource wastage. Combining them, M 2 NDP achieves significant speedups for various workloads by up to 128x (14.5x overall) and reduces energy by up to 87.9% (80.3% overall) compared to baseline CPU/GPU hosts with passive CXL memory. The M
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2de036f-d2b9-477e-8b86-a4be52b59c72Cited by top-tier papers4
- Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation MethodologyKonstantinos Kanellopoulos, Konstantinos Sgouras, F. Nisa Bostanci, Andreas Kosmas Kakolyris et al.ASPLOS 2025 · 8 citations
- RosenBridge: A Framework for Enabling Express I/O Paths Across the Virtualization BoundaryShi Qiu, Li Wang, Jianqin Yan, Ruofan Xiong et al.FAST 2026 · 1 citation
- AXLE: Coordinated Offloading with Asynchronous Back-Streaming in Computational Memory SystemsSuyeon Lee, Kangkyu Park, Kwangsik Shin, Ada GavrilovskaISCA 2026
- MAC: Metadata Acceleration for Sustainable Performance in Big-Data Systems with CXL DRAMDusol Lee, Yan Sun, Houxiang Ji, Vinit Gupta et al.OSDI 2026
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
Related papers
- Stream-Based Data Placement for Near-Data Processing with Extended MemoryYiwei Li, Boyu Tian, Yi Ren, Mingyu GaoMICRO 2024 · 5 citations
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong et al.MICRO 2025 · 3 citations
- Exploring Performance and Cost Optimization with ASIC-Based CXL MemoryYupeng Tang, Ping Zhou, Wenhui Zhang, Henry Hu et al.EuroSys 2024 · 40 citations
- Systematic CXL Memory Characterization and Performance Analysis at ScaleJinshu Liu, Hamid Hadian, Yuyue Wang, Daniel S. Berger et al.ASPLOS 2025 · 47 citations
- BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL SupportWenqin Huangfu, Krishna T. Malladi, Andrew Chang, Yuan XieMICRO 2022 · 24 citations
