ZnG: Architecting GPU Multi-Processors with New Flash for Scalable Data Analysis
Jie Zhang, Myoungsoo Jung
Abstract
We propose ZnG, a new GPU-SSD integrated architecture, which can maximize the memory capacity in a GPU and address performance penalties imposed by an SSD. Specifically, ZnG replaces all GPU internal DRAMs with an ultra-low-latency SSD to maximize the GPU memory capacity. ZnG further removes performance bottleneck of the SSD by replacing its flash channels with a high-throughput flash network and integrating SSD firmware in the GPU's MMU to reap the benefits of hardware accelerations. Although flash arrays within the SSD can deliver high accumulated bandwidth, only a small fraction of such bandwidth can be utilized by GPU's memory requests due to mismatches of their access granularity. To address this, ZnG employs a large L2 cache and flash registers to buffer the memory requests. Our evaluation results indicate that ZnG can achieve 7.5× higher performance than prior work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d0d3d1b-4249-45b6-8d4b-ce3ac4a5a0baCited by top-tier papers10
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min et al.ASPLOS 2023 · 48 citations
- Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNsShine Kim, Yunho Jin, Gina Sohn, Jonghyun Bae et al.FAST 2021 · 44 citations
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang et al.HPCA 2024 · 27 citations
- Bandwidth-Effective DRAM Cache for GPU s with Storage-Class MemoryJeongmin Hong, Sungjun Cho, Geonwoo Park, Wonhyuk Yang et al.HPCA 2024 · 21 citations
- G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor MigrationsHaoyang Zhang, Yirui Eric Zhou, Yuqi Xue, Yiqi Liu et al.MICRO 2023 · 21 citations
Related papers
- Ohm-GPU: Integrating New Optical Network and Heterogeneous Memory into GPU Multi-ProcessorsJie Zhang, Myoungsoo JungMICRO 2021 · 6 citations
- RayN: Ray Tracing Acceleration with Near-memory ComputingMohammadreza Saed, Prashant J. Nair, Tor M. AamodtMICRO 2025 · 5 citations
- Demand Layering for Real-Time DNN Inference with Minimized Memory UsageMingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn et al.RTSS 2022 · 21 citations
- Networked SSD: Flash Memory Interconnection Network for High-Bandwidth SSDJiho Kim, Seokwon Kang, Yongjun Park, John KimMICRO 2022 · 18 citations
- GOLAP: A GPU-in-Data-Path Architecture for High-Speed OLAPNils Boeschen, Tobias Ziegler, Carsten BinnigSIGMOD 2025 · 18 citations
