A ve : Guiding Agentic GPU Optimization Using Data-Flow Invariants
Haohui Mai, Xiaoyan Guo, Xiangyun Ding, Daifeng Li, Qiuchu Yu, Chenzhun Guo, Cong Wang, Jiacheng Zhao, Christos Kozyrakis, Binhang Yuan
Abstract
LLM coding agents can generate correct GPU kernels, yet their performance still lags behind expert libraries. Achieving peak GPU throughput requires coordinating low-level transformations such as shared-memory staging, software pipelining, and instruction scheduling. Unit tests and performance profiles provide only sparse end-to-end feedback, leaving agents unable to localize violations of the global constraints that these transformations must preserve.
We present Ave, an agentic framework that uses data-flow invariants as compile-time guardrails for kernel optimization. These invariants specify relationships among data values that must hold throughout execution. Ave's tile-based Pythonic DSL exposes hardware instructions and compiler policies while abstracting complex memory layout encodings as tiles. Tag functions assign symbolic labels derived from selected logical coordinates and expressions. The compiler propagates these labels through data and control flow, and tag assertions enforce required relationships among them at use sites. The compiler checks these assertions with a tractable, flow-sensitive, path-insensitive analysis backed by an SMT solver, returning concrete assignments that witness abstract violations and guide repairs with no runtime overhead. An in-context reinforcement learning (ICRL) planner proposes optimizations from a curated knowledge base. A lowering agent implements them and instantiates their invariants.
We evaluate Ave on AMD MI300X across GEMM, flash attention, and MoE, which together account for up to 90% of GPU time in LLM inference. With GPT-5.6 Sol, its optimized kernels achieve 89-99% of the effective throughput of stateof-the-art hand-optimized libraries and improve geometricmean throughput by 1.62-1176× over uncontaminated agentic baselines. On 200 KernelBench tasks, with GPT-5.6 Sol and in-context DSL examples, Ave produces valid kernels within
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdb63492-85c7-447d-8dd1-53bed60d9cb8Builds on9
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
- AKG: automatic kernel generation for neural processing units using polyhedral transformationsJie Zhao, Bojie Li, Wang Nie, Zhen Geng et al.PLDI 2021 · 81 citations
- Mirage: A Multi-Level Superoptimizer for Tensor ProgramsMengdi Wu, Xinhao Cheng, Shengyu Liu, Chunan Shi et al.OSDI 2025 · 49 citations
Related papers
- STARK: Strategic Team of Agents for Refining KernelsJuncheng Dong, Yang Yang, Tao Liu, Yang Wang et al.ICLR 2026 · 26 citations
- EGG: An Expert-Guided Agent Framework for Kernel GenerationYaochen Han, Ke Fan, Hongxu Jiang, Wanqi Xu et al.ICML 2026
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement LearningShiyang Li, Zijian Zhang, Winson Chen, Yuebo Luo et al.ICML 2026 · 9 citations
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel GenerationXinguo Zhu, Shaohui Peng, Jiaming Guo, Yunji Chen et al.AAAI 2026 · 9 citations
- Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel SynthesisYujie Zheng, Zhuo Li, Shengtao Zhang, Jiaqian Wang et al.ICML 2026 · 3 citations
