dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data Processing
Jiasheng Hu, Kaiwen Zheng, Anna Li, Sidharth Sankhe, Philip A. Bernstein, Qizhen Zhang
摘要
Data processing units, or DPUs, are equipped with hardware accelerators for compute-intensive data path tasks. Although DPUs' SoC cores are wimpier than the host's, hardware accelerators are typically orders of magnitude faster than CPUs. Harvesting DPU hardware accelerators for database systems could significantly increase throughput and save host CPU cycles. However, due to the heterogeneity of DPUs' hardware configurations and performance characteristics, it is challenging to offer a unified and portable solution for cloud data processing systems to harvest the compute resources on DPUs across generations and vendors. Additionally, due to DPU resource constraints, offloaded compute tasks need to be carefully optimized and scheduled to achieve high efficiency and avoid performance regression. To address these challenges, we introduce two levels of abstraction: dpKernels, which are unified, efficient, and portable primitives that abstract DPU compute resources (i.e., hardware accelerators and SoC cores) for cloud data systems, and dpManager, an onboard management framework that abstracts specific DPU platforms for dpKernels to deliver their promises with optimized, scheduled, and cross-platform executions. The benefits of our proposal have been validated by various workloads, systems, and DPU hardware.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu 等VLDB 2021 · 被引用 67 次
- Understanding the Effect of Data Center Resource Disaggregation on Production DBMSsQizhen Zhang, Yifan Cai, Xinyi Chen, Sebastian Angel 等VLDB 2020 · 被引用 64 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- Gimbal: enabling multi-tenant storage disaggregation on SmartNIC JBOFsJaehong Min, Ming Liu, Tapan Chugh, Chenxingyu Zhao 等SIGCOMM 2021 · 被引用 47 次
- The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar CodeAzim Afroozeh, Peter BonczVLDB 2023 · 被引用 44 次
相关 Paper
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu 等VLDB 2024 · 被引用 12 次
- Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUsDimitrios Giouroukis, Dwi P. A. Nugroho, Varun Pandey, Steffen Zeuch 等VLDB 2025 · 被引用 3 次
- DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data ProcessingChen Ding, Sicen Li, Kai Lu, Ting Yao 等USENIX ATC 2025 · 被引用 2 次
- NutCracker: A Compilation Framework for Hybrid DPU ArchitecturesYihan Yang, Haifeng Sun, Antoine Kaufmann, Jialin LiEuroSys 2026 · 被引用 2 次
- A Case for Graphics-driven Query ProcessingHarish Doraiswamy, Vikas Kalagi, Karthik Ramachandra, Jayant R. HaritsaVLDB 2023 · 被引用 4 次
