SC2021Top-tier venue
Dr. Top-k: delegate-centric Top-k on GPUs
Anil Gaihre, Da Zheng, Scott Weitze, Lingda Li, Shuaiwen Leon Song, Caiwen Ding, Xiaoye S. Li, Hang Liu
Abstract
Recent top-k computation efforts explore the possibility of revising various sorting algorithms to answer top-k queries on GPUs. These endeavors, unfortunately, perform significantly more work than needed. This paper introduces Dr. Top-k, a Delegate-centric top-k system on GPUs that can reduce the top-k workloads significantly. Particularly, it contains three major contributions: First, we introduce a comprehensive design of the delegate-centric concept, including maximum delegate, delegate-based filtering, and delegate mechanisms to help reduce the workload for top-k up to more than 99%. Second, due to the difficulty and importance of deriving a proper subrange size, we perform a rigorous theoretical analysis, coupled with thorough experimental validations to identify the desirable subrange size. Third, we introduce four key system optimizations to enable fast multi-GPU top-k computation. Taken together, this work constantly outperforms the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15aee1e2-a038-4331-bb58-13ff6d2a2ee7Cited by top-tier papers3
- GTS: GPU-based Tree Index for Fast Similarity SearchYifan Zhu, Ruiyao Ma, Baihua Zheng, Xiangyu Ke et al.SIGMOD 2024 · 8 citations
- BCCE: Block-Centric GPU Co-Design for Real-Time Range-Top-K Query at ScaleChengying Huan, Ziheng Meng, Zhengyi Yang, Yongchao Liu et al.HPDC 2026
- RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUsXi Xie, Yuebo Luo, Hongwu Peng, Caiwen DingICLR 2025
Builds on3
- PM-LSH: A Fast and Accurate LSH Framework for High-Dimensional Approximate NN SearchBolong Zheng, Xi Zhao, Lianggui Weng, Nguyen Quoc Viet Hung et al.VLDB 2020 · 64 citations
- FFT-based Gradient Sparsification for the Distributed Training of Deep Neural NetworksLinnan Wang, Wei Wu, Junyu Zhang, Hang Liu et al.HPDC 2020 · 21 citations
- Efficient Main-Memory Top-K Selection For Multicore ArchitecturesVasileios Zois, Vassilis J. Tsotras, Walid A. NajjarVLDB 2020 · 8 citations
Related papers
- Parallel Top-K Algorithms on GPU: A Comprehensive Study and New MethodsJingrong Zhang, Akira Naruse, Xipeng Li, Yong WangSC 2023 · 17 citations
- HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUsYanliang Li, Wenbo Li, Qian Gong, Qing Liu et al.SC 2025 · 2 citations
- Towards Energy-Efficient Real-Time Scheduling of Heterogeneous Multi-GPU SystemsYidi Wang, Mohsen Karimi, Hyoseung KimRTSS 2022 · 8 citations
- Realtime Top-k Personalized PageRank over Large Graphs on GPUsJieming Shi, Renchi Yang, Tianyuan Jin, Xiaokui Xiao et al.VLDB 2020 · 44 citations
- Near-optimal sparse allreduce for distributed deep learningShigang Li, Torsten HoeflerPPoPP 2022 · 57 citations
