A ve : Guiding Agentic GPU Optimization Using Data-Flow Invariants
Haohui Mai, Xiaoyan Guo, Xiangyun Ding, Daifeng Li, Qiuchu Yu, Chenzhun Guo, Cong Wang, Jiacheng Zhao, Christos Kozyrakis, Binhang Yuan
摘要
LLM coding agents can generate correct GPU kernels, yet their performance still lags behind expert libraries. Achieving peak GPU throughput requires coordinating low-level transformations such as shared-memory staging, software pipelining, and instruction scheduling. Unit tests and performance profiles provide only sparse end-to-end feedback, leaving agents unable to localize violations of the global constraints that these transformations must preserve.
We present Ave, an agentic framework that uses data-flow invariants as compile-time guardrails for kernel optimization. These invariants specify relationships among data values that must hold throughout execution. Ave's tile-based Pythonic DSL exposes hardware instructions and compiler policies while abstracting complex memory layout encodings as tiles. Tag functions assign symbolic labels derived from selected logical coordinates and expressions. The compiler propagates these labels through data and control flow, and tag assertions enforce required relationships among them at use sites. The compiler checks these assertions with a tractable, flow-sensitive, path-insensitive analysis backed by an SMT solver, returning concrete assignments that witness abstract violations and guide repairs with no runtime overhead. An in-context reinforcement learning (ICRL) planner proposes optimizations from a curated knowledge base. A lowering agent implements them and instantiates their invariants.
We evaluate Ave on AMD MI300X across GEMM, flash attention, and MoE, which together account for up to 90% of GPU time in LLM inference. With GPT-5.6 Sol, its optimized kernels achieve 89-99% of the effective throughput of stateof-the-art hand-optimized libraries and improve geometricmean throughput by 1.62-1176× over uncontaminated agentic baselines. On 200 KernelBench tasks, with GPT-5.6 Sol and in-context DSL examples, Ave produces valid kernels within
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- AKG: automatic kernel generation for neural processing units using polyhedral transformationsJie Zhao, Bojie Li, Wang Nie, Zhen Geng 等PLDI 2021 · 被引用 81 次
- Mirage: A Multi-Level Superoptimizer for Tensor ProgramsMengdi Wu, Xinhao Cheng, Shengyu Liu, Chunan Shi 等OSDI 2025 · 被引用 49 次
相关 Paper
- STARK: Strategic Team of Agents for Refining KernelsJuncheng Dong, Yang Yang, Tao Liu, Yang Wang 等ICLR 2026 · 被引用 26 次
- EGG: An Expert-Guided Agent Framework for Kernel GenerationYaochen Han, Ke Fan, Hongxu Jiang, Wanqi Xu 等ICML 2026
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement LearningShiyang Li, Zijian Zhang, Winson Chen, Yuebo Luo 等ICML 2026 · 被引用 9 次
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel GenerationXinguo Zhu, Shaohui Peng, Jiaming Guo, Yunji Chen 等AAAI 2026 · 被引用 9 次
- Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel SynthesisYujie Zheng, Zhuo Li, Shengtao Zhang, Jiaqian Wang 等ICML 2026 · 被引用 3 次
