Analytical characterization and design space exploration for optimization of CNNs
Rui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev, P. Sadayappan
Abstract
Moving data through the memory hierarchy is a fundamental bottleneck that can limit the performance of core algorithms of machine learning, such as convolutional neural networks (CNNs). Loop-level optimization, including loop tiling and loop permutation, are fundamental transformations to reduce data movement. However, the search space for finding the best loop-level optimization configuration is explosively large. This paper develops an analytical modeling approach for finding the best loop-level optimization configuration for CNNs on multi-core CPUs. Experimental evaluation shows that this approach achieves comparable or better performance than state-of-the-art libraries and auto-tuning based optimizers for CNNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bef0d07c-f2c2-4d67-a5de-fa0c532bcea2Cited by top-tier papers14
- ROLLER: Fast and Efficient Tensor Compilation for Deep LearningHongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke et al.OSDI 2022 · 84 citations
- LLMCompass: Enabling Efficient Hardware Design for Large Language Model InferenceHengrui Zhang, August Ning, Rohan Baskar Prabhakar, David WentzlaffISCA 2024 · 58 citations
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu et al.ASPLOS 2022 · 48 citations
- Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionSize Zheng, Siyuan Chen, Peidi Song, Renze Chen et al.HPCA 2023 · 46 citations
- DOSA: Differentiable Model-Based One-Loop Search for DNN AcceleratorsCharles Hong, Qijing Huang, Grace Dinh, Mahesh Subedar et al.MICRO 2023 · 19 citations
Related papers
- Optimizing the Memory Hierarchy by Compositing Automatic Transformations on Computations and DataJie Zhao, Peng DiMICRO 2020 · 32 citations
- I/O lower bounds for auto-tuning of convolutions in CNNsXiaoyang Zhang, Junmin Xiao, Guangming TanPPoPP 2021 · 11 citations
- IOOpt: automatic derivation of I/O complexity bounds for affine programsAuguste Olivry, Guillaume Iooss, Nicolas Tollenaere, Atanas Rountev et al.PLDI 2021 · 10 citations
- ALT: Breaking the Wall between Data Layout and Loop Optimizations for Deep Learning CompilationZhiying Xu, Jiafan Xu, Hongding Peng, Wei Wang et al.EuroSys 2023 · 12 citations
- Welder: Scheduling Deep Learning Memory Access via Tile-graphYining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma et al.OSDI 2023 · 64 citations
