Lune

ASPLOS2026Top-tier venue

LOOPRAG: Enhancing Loop Transformation Optimization with Retrieval-Augmented Large Language Models

Yijie Zhi, Yayu Cao, Jianhua Dai, Xiaoyang Han, Jingwen Pu, Qinran Wu, Sheng Cheng, Ming Cai

2026Year
1Citations

Abstract

Loop transformations are semantics-preserving optimization techniques applied at the source level, widely used in compilers to maximize objectives such as vectorization and parallelism. Despite decades of research, applying the optimal composition of loop transformations remains challenging due to inherent complexities, including dependency analysis and cost modeling for optimization objectives.

Recent studies have explored the potential of Large Language Models (LLMs) for code optimization. However, our key observation is that LLMs often struggle with effective loop transformation optimization, frequently leading to errors or suboptimal optimization, thereby missing significant opportunities for performance improvements.

To bridge this gap, we propose LOOPRAG, a novel retrieval-augmented generation framework designed to guide LLMs in performing effective loop optimization on Static Control Part (SCoP). We introduce a parameter-driven method to harness loop properties, which trigger various loop transformations, and generate diverse yet legal example codes serving as a demonstration source. To effectively obtain the most informative demonstrations, we propose a loop-aware algorithm based on loop features, which balances similarity and diversity for code retrieval. To enhance correct and efficient code generation, we introduce a feedback-based iterative mechanism that incorporates compilation, testing and performance results as feedback to guide LLMs. Each optimized code generated by LOOPRAG undergoes mutation, coverage and differential testing to perform equivalence checking.

We evaluate LOOPRAG on PolyBench, TSVC and LORE benchmark suites, and compare it against compilers (GCC-Graphite, Clang-Polly, Perspective and ICX) and representative LLMs (DeepSeek and GPT-4). The results demonstrate average speedups over compilers of up to 11.20×, 14.34×, and 9.29× for PolyBench, TSVC, and LORE, respectively, and speedups over base LLMs of up to 11.97×, 5.61×, and 11.59×.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 1bd2b663-89ab-45a3-bb3d-af93b557a930

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines