USENIX ATC2025顶会
DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data Processing
Chen Ding, Sicen Li, Kai Lu, Ting Yao, Daohui Wang, Huatao Wu, Jiguang Wan, Zhihu Tan, Changsheng Xie
摘要
Shuffle is a crucial operation in distributed data processing, responsible for transferring intermediate data between nodes. However, it is highly resource-intensive, consuming significant CPU power and often becoming a major performance bottleneck, particularly in data analysis tasks involving large datasets.
In this paper, we introduce DShuffle, an efficient framework that leverages DPUs to offload and accelerate shuffle operations. The DPU, with its specialized compute and I/O hardware, is ideally suited for offloading on-path shuffle tasks. However, its complex architecture requires careful design for effective offloading. To fully harness the DPU's capabilities, DShuffle divides the shuffle process into three stages: serialization, preprocessing, and I/O, and organizes them in a pipelined manner for efficient execution on the DPU. By leveraging high-concurrency memory access units to accelerate the serialization phase and using the DPU to directly write intermediate data to disk, DShuffle effectively accelerates the shuffle process and eliminates unnecessary data copies. Our experiments on a real DPU platform with industrial-grade Spark demonstrate that DShuffle enhances both host CPU and I/O efficiency and effectively reduce Spark task completion times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im 等SOSP 2021 · 被引用 83 次
- Xenic: SmartNIC-Accelerated Distributed TransactionsHenry N. Schuh, Weihao Liang, Ming Liu, Jacob Nelson 等SOSP 2021 · 被引用 62 次
- Gimbal: enabling multi-tenant storage disaggregation on SmartNIC JBOFsJaehong Min, Ming Liu, Tapan Chugh, Chenxingyu Zhao 等SIGCOMM 2021 · 被引用 47 次
- A Specialized Architecture for Object Serialization with Applications to Big Data AnalyticsJaeyoung Jang, Sungjun Jung, Sunmin Jeong, Jun Heo 等ISCA 2020 · 被引用 32 次
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 被引用 24 次
相关 Paper
- dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data ProcessingJiasheng Hu, Kaiwen Zheng, Anna Li, Sidharth Sankhe 等VLDB 2026
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu 等VLDB 2024 · 被引用 12 次
- OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object StorageZhen Jin, Yiquan Chen, Mingxu Liang, Yijing Wang 等ASPLOS 2025 · 被引用 5 次
- DFlush: DPU-Offloaded Flush for Disaggregated LSM-based Key-Value StoresChen Ding, Kai Lu, Quanyi Zhang, Zekun Ye 等SIGMOD 2025 · 被引用 7 次
- MinFlow: High-performance and Cost-efficient Data Passing for I/O-intensive Stateful Serverless AnalyticsTao Li, Yongkun Li, Wenzhe Zhu, Yinlong Xu 等FAST 2024 · 被引用 5 次
