DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data Processing
Chen Ding, Sicen Li, Kai Lu, Ting Yao, Daohui Wang, Huatao Wu, Jiguang Wan, Zhihu Tan, Changsheng Xie
Abstract
Shuffle is a crucial operation in distributed data processing, responsible for transferring intermediate data between nodes. However, it is highly resource-intensive, consuming significant CPU power and often becoming a major performance bottleneck, particularly in data analysis tasks involving large datasets.
In this paper, we introduce DShuffle, an efficient framework that leverages DPUs to offload and accelerate shuffle operations. The DPU, with its specialized compute and I/O hardware, is ideally suited for offloading on-path shuffle tasks. However, its complex architecture requires careful design for effective offloading. To fully harness the DPU's capabilities, DShuffle divides the shuffle process into three stages: serialization, preprocessing, and I/O, and organizes them in a pipelined manner for efficient execution on the DPU. By leveraging high-concurrency memory access units to accelerate the serialization phase and using the DPU to directly write intermediate data to disk, DShuffle effectively accelerates the shuffle process and eliminates unnecessary data copies. Our experiments on a real DPU platform with industrial-grade Spark demonstrate that DShuffle enhances both host CPU and I/O efficiency and effectively reduce Spark task completion times.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0372e184-41bc-44c5-9ee5-845560ef9ccdBuilds on13
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im et al.SOSP 2021 · 83 citations
- Xenic: SmartNIC-Accelerated Distributed TransactionsHenry N. Schuh, Weihao Liang, Ming Liu, Jacob Nelson et al.SOSP 2021 · 62 citations
- Gimbal: enabling multi-tenant storage disaggregation on SmartNIC JBOFsJaehong Min, Ming Liu, Tapan Chugh, Chenxingyu Zhao et al.SIGCOMM 2021 · 47 citations
- A Specialized Architecture for Object Serialization with Applications to Big Data AnalyticsJaeyoung Jang, Sungjun Jung, Sunmin Jeong, Jun Heo et al.ISCA 2020 · 32 citations
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 24 citations
Related papers
- dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data ProcessingJiasheng Hu, Kaiwen Zheng, Anna Li, Sidharth Sankhe et al.VLDB 2026
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu et al.VLDB 2024 · 12 citations
- OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object StorageZhen Jin, Yiquan Chen, Mingxu Liang, Yijing Wang et al.ASPLOS 2025 · 5 citations
- DFlush: DPU-Offloaded Flush for Disaggregated LSM-based Key-Value StoresChen Ding, Kai Lu, Quanyi Zhang, Zekun Ye et al.SIGMOD 2025 · 7 citations
- MinFlow: High-performance and Cost-efficient Data Passing for I/O-intensive Stateful Serverless AnalyticsTao Li, Yongkun Li, Wenzhe Zhu, Yinlong Xu et al.FAST 2024 · 5 citations
