A Framework for Fine-Grained Program Versioning
Yishen Chen, Saman P. Amarasinghe
Abstract
Static dependence analysis is critical for optimizations such as vectorization and loop-invariant code motion. However, traditional static dependence analysis is often imprecise, making these optimizations less effective. To address this issue, production compilers use loop versioning to rule out some categories of memory dependencies at run time. However, loop versioning is loop-centric and usually tied to specific optimizations (e.g., loop vectorization), making it less effective for nonloop optimizations such as superword-level parallelism (SLP) vectorization.
In this paper, we propose a fine-grained versioning framework to rule out program dependencies at run time. Our framework is general and not tailored to any specific optimizations. To use our system, a client optimization specifies groups of instructions (or loops) whose independence is desired but unprovable statically. In response, our system duplicates the appropriate instructions and guards the original ones with run-time checks to guarantee their independence; if the checks fail, the duplicated instructions execute instead.
In a case study, we extended an existing SLP vectorizer with minimal modifications using our framework, resulting in a 1.17× speedup over Clang's vectorizers on TSVC and a 1.51× speedup on PolyBench. In both benchmarks, we encountered programs that could not be vectorized with loop versioning alone.
In a second case study, we used our framework to implement a more aggressive variant of redundant load elimination than the one implemented by Clang. Our redundant load elimination results in a 1.012× speedup on the SPEC 2017 Floating Point benchmarks, with the maximum speedup being 1.064×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d6a5676-43aa-4de4-8d49-05c2de7b1870Builds on3
- All you need is superword-level parallelism: systematic control-flow vectorization with SLPYishen Chen, Charith Mendis, Saman P. AmarasinghePLDI 2022 · 20 citations
- The road not taken: exploring alias analysis based optimizations missed by the compilerKhushboo Chitre, Piyus Kedia, Rahul PurandareOOPSLA 2022 · 13 citations
- Rapid: Region-Based Pointer DisambiguationKhushboo Chitre, Piyus Kedia, Rahul PurandareOOPSLA 2023 · 2 citations
Related papers
- Speculative Vectorisation with Selective ReplayPeng Sun, Giacomo Gabrielli, Timothy M. JonesISCA 2021 · 3 citations
- SCAF: a speculation-aware collaborative dependence analysis frameworkSotiris Apostolakis, Ziyang Xu, Zujun Tan, Greg Chan et al.PLDI 2020 · 12 citations
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones et al.MICRO 2023 · 15 citations
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 27 citations
- An abstract interpretation for SPMD divergence on reducible control flow graphsJulian Rosemann, Simon Moll, Sebastian HackPOPL 2021 · 8 citations
