Hardware/Software Co-Programmable Framework for Computational SSDs to Accelerate Deep Learning Service on Large-Scale Graphs
Miryeong Kwon, Donghyun Gouk, Sangwon Lee, Myoungsoo Jung
摘要
Graph neural networks (GNNs) process large-scale graphs consisting of a hundred billion edges. In contrast to traditional deep learning, unique behaviors of the emerging GNNs are engaged with a large set of graphs and embedding data on storage, which exhibits complex and irregular preprocessing.
We propose a novel deep learning framework on large graphs, HolisticGNN, that provides an easy-to-use, nearstorage inference infrastructure for fast, energy-efficient GNN processing. To achieve the best end-to-end latency and high energy efficiency, HolisticGNN allows users to implement various GNN algorithms and directly executes them where the actual data exist in a holistic manner. It also enables RPC over PCIe such that the users can simply program GNNs through a graph semantic library without any knowledge of the underlying hardware or storage configurations.
We fabricate HolisticGNN's hardware RTL and implement its software on an FPGA-based computational SSD (CSSD). Our empirical evaluations show that the inference time of HolisticGNN outperforms GNN inference services using high-performance modern GPUs by 7.1× while reducing energy consumption by 33.2×, on average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- CXL-ANNS: Software-Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor SearchJunhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Seungjun Lee 等USENIX ATC 2023 · 被引用 75 次
- NVMeVirt: A Versatile Software-defined Virtual NVMe DeviceSang-Hoon Kim, Jaehoon Shim, Euidong Lee, Seong-Yeob Jeong 等FAST 2023 · 被引用 52 次
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang 等HPCA 2024 · 被引用 27 次
- In-Storage Domain-Specific Acceleration for Serverless ComputingRohan Mahapatra, Soroush Ghodrati, Byung Hoon Ahn, Sean Kinzer 等ASPLOS 2024 · 被引用 7 次
- ScalaCache: Scalable User-Space Page Cache Management with Software-Hardware CoordinationLi Peng, Yuda An, You Zhou, Chenxi Wang 等USENIX ATC 2024 · 被引用 6 次
它引用的顶会 Paper6
- HyGCN: A GCN Accelerator with Hybrid ArchitectureMingyu Yan, Lei Deng, Xing Hu, Ling Liang 等HPCA 2020 · 被引用 338 次
- Hardware Acceleration of Graph Neural NetworksAdam Auten, Matthew Tomei, Rakesh KumarDAC 2020 · 被引用 108 次
- RecSSD: near data processing for solid state drive based recommendation inferenceMark Wilkening, Udit Gupta, Samuel Hsia, Caroline Trippel 等ASPLOS 2021 · 被引用 100 次
- Accelerating graph sampling for graph machine learning using GPUsAbhinav Jangda, Sandeep Polisetty, Arjun Guha, Marco SerafiniEuroSys 2021 · 被引用 79 次
- A Fast and Flexible Hardware-based Virtualization Mechanism for Computational Storage DevicesDongup Kwon, Dongryeong Kim, Junehyuk Boo, Wonsik Lee 等USENIX ATC 2021 · 被引用 22 次
相关 Paper
- WholeGraph: A Fast Graph Neural Network Training Framework with Multi-GPU Distributed Shared Memory ArchitectureDongxu Yang, Junhong Liu, Jiaxing Qi, Junjie LaiSC 2022 · 被引用 12 次
- SmartSAGE: training large-scale graph neural networks using in-storage processing architecturesYunjae Lee, Jinha Chung, Minsoo RhuISCA 2022 · 被引用 57 次
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 被引用 37 次
- FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleZeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li 等ASPLOS 2024 · 被引用 8 次
- Hyperscale FPGA-as-a-service architecture for large-scale distributed graph neural networkShuangchen Li, Dimin Niu, Yuhao Wang, Wei Han 等ISCA 2022 · 被引用 30 次
