Efficient automatic scheduling of imaging and vision pipelines for the GPU
Luke Anderson, Andrew Adams, Karima Ma, Tzu-Mao Li, Tian Jin, Jonathan Ragan-Kelley
摘要
We present a new algorithm to quickly generate high-performance GPU implementations of complex imaging and vision pipelines, directly from high-level Halide algorithm code. It is fully automatic, requiring no schedule templates or hand-optimized kernels. We address the scalability challenge of extending search-based automatic scheduling to map large real-world programs to the deep hierarchies of memory and parallelism on GPU architectures in reasonable compile time. We achieve this using (1) a two-phase search algorithm that first 'freezes' decisions for the lowest cost sections of a program, allowing relatively more time to be spent on the important stages, (2) a hierarchical sampling strategy that groups schedules based on their structural similarity, then samples representatives to be evaluated, allowing us to explore a large space with few samples, and (3) memoization of repeated partial schedules, amortizing their cost over all their occurrences. We guide the process with an efficient cost model combining machine learning, program analysis, and GPU architecture knowledge.
We evaluate our method's performance on a diverse suite of real-world imaging and vision pipelines. Our scalability optimizations lead to average compile time speedups of 49× (up to 530×). We find schedules that are on average 1.7× faster than existing automatic solutions (up to 5×), and competitive with what the best human experts were able to achieve in an active effort to beat our automatic results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- TLP: A Deep Learning-Based Cost Model for Tensor Program TuningYi Zhai, Yu Zhang, Shuo Liu, Xiaomeng Chu 等ASPLOS 2023 · 被引用 42 次
- Autoscheduling for sparse tensor algebra with an asymptotic cost modelWillow Ahrens, Fredrik Kjolstad, Saman P. AmarasinghePLDI 2022 · 被引用 30 次
- Mosaic: An Interoperable Compiler for Tensor AlgebraManya Bansal, Olivia Hsu, Kunle Olukotun, Fredrik KjolstadPLDI 2023 · 被引用 16 次
- Optimal Kernel Orchestration for Tensor Programs with KorchMuyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu 等ASPLOS 2024 · 被引用 11 次
- Guided Equality SaturationThomas Koehler, Andrés Goens, Siddharth Bhat, Tobias Grosser 等POPL 2024 · 被引用 9 次
它引用的顶会 Paper5
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen 等ASPLOS 2020 · 被引用 171 次
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network CompilationByung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, Hadi EsmaeilzadehICLR 2020 · 被引用 90 次
- Reinforced Genetic Algorithm Learning for Optimizing Computation GraphsAditya Paliwal, Felix Gimeno, Vinod Nair, Yujia Li 等ICLR 2020 · 被引用 70 次
- Transferable Graph Optimizers for ML CompilersYanqi Zhou, Sudip Roy, AmirAli Abdolrashidi, Daniel Wong 等NeurIPS 2020 · 被引用 63 次
相关 Paper
- The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSLBo Qiao, M. Akif Özkan, Jürgen Teich, Frank HannigDAC 2020 · 被引用 8 次
- Aδ: autodiff for discontinuous programs - applied to shadersYuting Yang, Connelly Barnes, Andrew Adams, Adam FinkelsteinSIGGRAPH 2022 · 被引用 17 次
- MISAAL: Synthesis-Based Automatic Generation of Efficient and Retargetable Semantics-Driven OptimizationsAbdul Rafae Noor, Dhruv Baronia, Akash Kothari, Muchen Xu 等PLDI 2025 · 被引用 2 次
- Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware ArchitecturesAkash Kothari, Abdul Rafae Noor, Muchen Xu, Hassam Uddin 等ASPLOS 2024 · 被引用 8 次
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUsRupanshu Soi, Rohan Yadav, Fredrik Kjolstad, Alex Aiken 等OSDI 2026 · 被引用 9 次
