PIMDup: An Optimized Deduplication Design on a Real Processing-in-Memory System
Chun-Le Yeh, Liang-Chi Chen, Chien-Chung Ho, Yu-Ming Chang, Da-Wei Chang
Abstract
Data deduplication enhances storage efficiency through non-destructive compression but is often hindered by the chunking process, which requires scanning the entire dataset. While traditional methods leveraging conventional architectures and hardware accelerators (e.g., GPUs and FPGAs) have been developed to address this issue, they continue to face challenges related to excessive data movement and associated performance degradation. These limitations stem from the von Neumann architecture, where computation and storage are separated in a processor-centric design, necessitating multiple memory hierarchy traversals and causing inefficiencies. To overcome these challenges, we explore UPMEM’s DPU, a processing-in-memory (PIM) technology that reduces data movement by performing computations directly within memory. However, designing a deduplication system for DPUs presents unique obstacles, including restricted inter-DPU data sharing, the absence of native multiplication support, and significant DPU-CPU communication overhead. In response, we propose PIMDup, a DPU-optimized deduplication system that addresses these constraints through efficient parallelization, DPU-friendly chunking techniques, and reduced data transfer volumes. Experimental results demonstrate that PIMDup improves chunking performance without compromising deduplication accuracy, achieving a speedup over CPU-based systems while maintaining 100% result consistency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1a7995fb-5ba8-4a4d-805c-2cf8e95dcea7Related papers
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho et al.HPCA 2025 · 7 citations
- PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM SystemsDongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo RhuMICRO 2024 · 23 citations
- PIMLex: A High-Performance Learned Index with Processing-in-MemoryLixiao Cui, Kedi Yang, Yusen Li, Gang Wang et al.FAST 2025 · 11 citations
- Accelerating Transactional Execution via Processing-In-MemoryAndré Lopes, Daniel Castro, Paolo RomanoEuroSys 2026
- Accelerating Aggregation Using a Real Processing-in-Memory SystemMuhammad Attahir Jibril, Hani Al-Sayeh, Kai-Uwe SattlerICDE 2024 · 7 citations
