SC2020Top-tier venue
Compiling generalized histograms for GPU
Troels Henriksen, Sune Hellfritzsch, Ponnuswamy Sadayappan, Cosmin E. Oancea
Abstract
We present and evaluate an implementation technique for histogram-like computations on GPUs that ensures both work-efficient asymptotic cost, support for arbitrary associative and commutative operators, and efficient use of hardware-supported atomic operations when applicable. Based on a systematic empirical examination of the design space, we develop a technique that balances conflict rates and memory footprint. We demonstrate our technique both as a library implementation in CUDA, as well as by extending the parallel array language Futhark with a new construct for expressing generalized histograms, and by supporting this construct with several compiler optimizations. We show that our histogram implementation taken in isolation outperforms similar primitives from CUB, and that it is competitive or outperforms the hand-written code of several application benchmarks, even when the latter is specialized for a class of datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c0751625-3d1c-449f-9c88-71489e153acbCited by top-tier papers2
- AD for an Array Language with Nested ParallelismRobert Schenck, Ola Rønning, Troels Henriksen, Cosmin E. OanceaSC 2022 · 11 citations
- Verifying Array Properties in Pure Data-Parallel ProgramsNikolaj Hey Hinnerskov, Robert Schenck, Cosmin E. OanceaPLDI 2026 · 1 citation
Related papers
- Descend: A Safe GPU Systems Programming LanguageBastian Köpcke, Sergei Gorlatch, Michel SteuwerPLDI 2024 · 7 citations
- AUTOMAP: Inferring Rank-Polymorphic Function Applications with Integer Linear ProgrammingRobert Schenck, Nikolaj Hey Hinnerskov, Troels Henriksen, Magnus Madsen et al.OOPSLA 2024 · 1 citation
- CUDASTF: Bridging the Gap Between CUDA and Task ParallelismCédric Augonnet, Andrei Alexandrescu, Albert Sidelnik, Michael GarlandSC 2024 · 7 citations
- Kuiper: Correct and Efficient GPU Programming with Dependent Types and Separation LogicGuido Martínez, Bastian Köpcke, Jonás Fiala, Gabriel Ebner et al.PLDI 2026 · 2 citations
- The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSLBo Qiao, M. Akif Özkan, Jürgen Teich, Frank HannigDAC 2020 · 8 citations
