(MC)2: Lazy MemCopy at the Memory Controller
Aditya K. Kamath, Simon Peter
Abstract
(MC)2is a lazy memory copy mechanism which can be used within memcpy-like functions to significantly reduce the CPU overhead for copies that are sparsely accessed. It can also hide copy latencies by enhancing the CPU’s ability to execute them asynchronously. (MC)2,s lazy memcpy avoids copying data at the time of invocation. Instead, (MC)2tracks prospective copies. If copied data is later accessed by a CPU or the cache, (MC)2uses the tracking information to lazily execute a copy, when necessary. Placing (MC)2at the memory controller puts it at the perfect vantage point to eliminate the largest source of memcpy overhead–CPU stalls due to cache misses in the critical path–while imposing minimal overhead itself. (MC)2consists of three main components: memory controller extensions that implement a lazy memcpy operation, a new instruction exposing the lazy memcpy, and a flexible software wrapper with semantics identical to memcpy. We implement and evaluate (MC)2in the gem5 simulator using a variety of microbenchmarks and workloads, including Google’s Protobuf, where (MC)2provides a speedup and Linux huge page copy-on-write faults, where (MC)2provides lower latency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a2d725b3-62cd-42ff-aa05-a541eda57524Cited by top-tier papers1
Ask how each one uses itRelated papers
- The Sparsity-Aware LazyGPU ArchitectureChangxi Liu, Miao Yu, Yifan Sun, Trevor E. CarlsonISCA 2025 · 1 citation
- How to Copy Memory? Coordinated Asynchronous Copy as a First-Class OS ServiceJingkai He, Yunpeng Dong, Dong Du, Mo Zou et al.SOSP 2025 · 2 citations
- zIO: Accelerating IO-Intensive Applications with Transparent Zero-Copy IOTimothy Stamler, Deukyeon Hwang, Amanda Raybuck, Wei Zhang et al.OSDI 2022 · 21 citations
- GhOST: a GPU Out-of-Order Scheduling Technique for Stall ReductionIshita Chaturvedi, Bhargav Reddy Godala, Yucan Wu, Ziyang Xu et al.ISCA 2024 · 9 citations
- RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency HidingNingzhi Ai, Wenjian He, Hu He, Jing Xia et al.MICRO 2025 · 2 citations
