OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object Storage
Zhen Jin, Yiquan Chen, Mingxu Liang, Yijing Wang, Guoju Fang, Ao Zhou, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenkai Shi
摘要
Object storage is increasingly attractive for deep learning (DL) applications due to its cost-effectiveness and high scalability. However, it exacerbates CPU burdens in DL clusters due to intensive object storage processing and multiple data movements. Data processing unit (DPU) offloading is a promising solution, but naively offloading the existing object storage client leads to severe performance degradation. Besides, only offloading the object storage client still involves redundant data movements, as data must first transfer from the DPU to the host and then from the host to the GPU, which continues to consume valuable host resources.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding PredictionGuofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu 等ICDE 2026
- CETOFS: A High-Performance File System with Host-Server Collaboration for Remote StorageWenqing Jia, Dejun Jiang, Jin XiongFAST 2026
- RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICsMaximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu 等OSDI 2026
- SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNICXuzheng Chen, Jie Zhang, Baolin Zhu, Xueying Zhu 等EuroSys 2026
相关 Paper
- SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersHanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang 等EuroSys 2023 · 被引用 22 次
- Managing Scalable Direct Storage Accesses for GPUs with GoFSShaobo Li, Yirui Eric Zhou, Yuqi Xue, Yuan Xu 等SOSP 2025
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu 等VLDB 2024 · 被引用 12 次
- DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data ProcessingChen Ding, Sicen Li, Kai Lu, Ting Yao 等USENIX ATC 2025 · 被引用 2 次
- NDPipe: Exploiting Near-data Processing for Scalable Inference and Continuous Training in Photo StorageJungwoo Kim, Seonggyun Oh, Jaeha Kung, Yeseong Kim 等ASPLOS 2024
