SmartDIMM: In-Memory Acceleration of Upper Layer Protocols
Neel Patel, Amin Mamandipoor, Mohammad Nouri, Mohammad Alian
Abstract
There has been significant focus on offloading upperlayer network protocols (ULPs) to accelerators located on CPUs and SmartNICs. However, restricting accelerator placement to these locations limits both the variety of ULPs that can be accelerated and the overall performance. In particular, it overlooks the opportunity to accelerate ULPs running atop a stateful transport protocol in the face of high cache contention. That is, at high network rates, the frequent DRAM accesses and SmartNIC-CPU synchronizations outweigh the benefits of hardware acceleration. This work introduces SmartDIMM, which unlocks the opportunity for accelerating ULPs running atop stateful transport protocols that primarily operate on data stored in DRAM. We prototyped SmartDIMM using Samsung's AxDIMM and implemented endto-end offloading of (de/en)cryption and (de)compression– two ULPs widely employed in datacenters. We then compared the performance of SmartDIMM with accelerator placements on the CPU, SmartNIC, and PCIe cards. Our results demonstrate that ULP offloading on SmartDIMM outperforms CPU, SmartNIC and PCIe-based offload configurations. In comparison to a server executing (de/en)cryption and (de)compression on the CPU, SmartDIMM achieves 21.0% to 10.28 × higher requests per second and 36.3% to 88.9% lower memory bandwidth utilization.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d7167eb0-5991-4809-9b42-2798c62b5efeCited by top-tier papers2
- Accelerating Retrieval-Augmented GenerationDerrick Quinn, Mohammad Nouri, Neel Patel, John Salihu et al.ASPLOS 2025 · 37 citations
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim et al.HPCA 2025 · 5 citations
Related papers
- SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNICXuzheng Chen, Jie Zhang, Baolin Zhu, Xueying Zhu et al.EuroSys 2026
- BASK: Batch And SmartNIC-offloaded KSMChanshin Kwak, Jaehyeon Lee, Minkyu Jung, Changjun Lee et al.EuroSys 2026
- STYX: Exploiting SmartNIC Capability to Reduce Datacenter Memory TaxHouxiang Ji, Mark Mansi, Yan Sun, Yifan Yuan et al.USENIX ATC 2023 · 22 citations
- Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUsDimitrios Giouroukis, Dwi P. A. Nugroho, Varun Pandey, Steffen Zeuch et al.VLDB 2025 · 3 citations
- AsyncDIMM: Achieving Asynchronous Execution in DIMM-Based Near-Memory ProcessingLiyan Chen, Dongxu Lyu, Jianfei Jiang, Qin Wang et al.HPCA 2025 · 7 citations
