RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
Xi Xie, Yuebo Luo, Hongwu Peng, Caiwen Ding
摘要
Top-k selection algorithms are fundamental in a wide range of applications, including high-performance computing, information retrieval, big data processing, and neural network model training. In this paper, we present RTop-K, a highly efficient parallel row-wise top-k selection algorithm specifically designed for GPUs. RTop-K leverages a binary search-based approach to optimize row-wise top-k selection, providing a scalable and accelerated solution. We conduct a detailed analysis of early stopping in our algorithm, showing that it effectively maintains the testing accuracy of neural network models while substantially improving performance. Our GPU implementation of RTop-K demonstrates superior performance over state-of-the-art row-wise top-k GPU implementations, achieving an average speed-up of up to 11.49× with early stopping and 7.29× without early stopping. Moreover, RTop-K accelerates the overall training workflow of MaxK-GNNs, delivering speed-ups ranging from 11.97% to 33.29% across different models and datasets. The GPU implementation can be found on Github † .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SonicMoE: Accelerating MoE with IO and Tile-aware OptimizationsWentao Guo, Mayank Mishra, Xinle Cheng, Ion Stoica 等ICLR 2026 · 被引用 22 次
- Stochastic Sparse Attention for Memory-Bound InferenceKyle Lee, Corentin Delacour, Kevin Callahan-Coray, Kyle Jiang 等ICML 2026
它引用的顶会 Paper6
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan 等ICLR 2020 · 被引用 1,155 次
- Top-KAST: Top-K Always Sparse TrainingSiddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero 等NeurIPS 2020 · 被引用 116 次
- MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks TrainingHongwu Peng, Xi Xie, Kaustubh Shivdikar, Md Amit Hasan 等ASPLOS 2024 · 被引用 32 次
- Parallel Top-K Algorithms on GPU: A Comprehensive Study and New MethodsJingrong Zhang, Akira Naruse, Xipeng Li, Yong WangSC 2023 · 被引用 17 次
相关 Paper
- Dr. Top-k: delegate-centric Top-k on GPUsAnil Gaihre, Da Zheng, Scott Weitze, Lingda Li 等SC 2021 · 被引用 16 次
- RTNN: accelerating neighbor search using hardware ray tracingYuhao ZhuPPoPP 2022 · 被引用 43 次
- PruneGNN: Algorithm-Architecture Pruning Framework for Graph Neural Network AccelerationDeniz Gurevin, Mohsin Shan, Shaoyi Huang, Md Amit Hasan 等HPCA 2024 · 被引用 28 次
- BCCE: Block-Centric GPU Co-Design for Real-Time Range-Top-K Query at ScaleChengying Huan, Ziheng Meng, Zhengyi Yang, Yongchao Liu 等HPDC 2026
- PilotANN: Memory-Bounded GPU Acceleration for Vector SearchYuntao Gui, Peiqi Yin, Xiao Yan, Chaorui Zhang 等KDD 2026 · 被引用 5 次
