Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders
Hyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Euicheol Lim, Gwangsun Kim
摘要
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not costeffective for NDP because they are not optimized for memorybound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur µs-scale latency and are not suitable for fine-grained NDP.
To achieve high-performance NDP end-to-end, we propose a low-overhead general-purpose NDP architecture for CXL memory referred to as Memory-Mapped NDP (M 2 NDP), which comprises memory-mapped functions (M 2 func) and memory-mapped µthreading (M 2 µthread). M 2 func is a CXL.mem-compatible lowoverhead communication mechanism between the host processor and NDP controller in CXL memory. M 2 µthread enables lowcost, general-purpose NDP unit design by introducing lightweight µthreads that support highly concurrent execution of kernels with minimal resource wastage. Combining them, M 2 NDP achieves significant speedups for various workloads by up to 128x (14.5x overall) and reduces energy by up to 87.9% (80.3% overall) compared to baseline CPU/GPU hosts with passive CXL memory. The M
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation MethodologyKonstantinos Kanellopoulos, Konstantinos Sgouras, F. Nisa Bostanci, Andreas Kosmas Kakolyris 等ASPLOS 2025 · 被引用 8 次
- RosenBridge: A Framework for Enabling Express I/O Paths Across the Virtualization BoundaryShi Qiu, Li Wang, Jianqin Yan, Ruofan Xiong 等FAST 2026 · 被引用 1 次
- AXLE: Coordinated Offloading with Asynchronous Back-Streaming in Computational Memory SystemsSuyeon Lee, Kangkyu Park, Kwangsik Shin, Ada GavrilovskaISCA 2026
- MAC: Metadata Acceleration for Sustainable Performance in Big-Data Systems with CXL DRAMDusol Lee, Yan Sun, Houxiang Ji, Vinit Gupta 等OSDI 2026
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah 等ISCA 2024 · 被引用 282 次
相关 Paper
- Stream-Based Data Placement for Near-Data Processing with Extended MemoryYiwei Li, Boyu Tian, Yi Ren, Mingyu GaoMICRO 2024 · 被引用 5 次
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong 等MICRO 2025 · 被引用 3 次
- Exploring Performance and Cost Optimization with ASIC-Based CXL MemoryYupeng Tang, Ping Zhou, Wenhui Zhang, Henry Hu 等EuroSys 2024 · 被引用 40 次
- Systematic CXL Memory Characterization and Performance Analysis at ScaleJinshu Liu, Hamid Hadian, Yuyue Wang, Daniel S. Berger 等ASPLOS 2025 · 被引用 47 次
- BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL SupportWenqin Huangfu, Krishna T. Malladi, Andrew Chang, Yuan XieMICRO 2022 · 被引用 24 次
