QiMeng-Tensify: Scaling Up Tensor Computation Optimization via Architecture-Aware LLM-Guided MCTS
Shouyang Dong, Jun Bi, Yuanbo Wen, Xiyue Yu, Jianxing Xu, Guanglin Xu, Ling Li, Xuehai Zhou, Tianshi Chen, Qi Guo
Abstract
The growing scale and complexity of large language models (LLMs) have intensified the need for optimizing largescale tensor computations (e.g., self-attention and mixture-ofexperts) on hardware platforms. Existing solutions rely on either manual expert optimization or exploration-based autotuning methods. However, neither approach scales effectively for LLMs with hundreds, even thousands of operators and dynamic control flows, because of prohibitive optimization overheads or suboptimal performance. To address this problem, we present QiMeng-Tensify, the first framework that combines LLMs with sequential decision optimization for large-scale graph-level tensor computation. Our key insight is that: (1) tensor computation optimization can be formulated as a generalized sequential decision problem to enlarge the optimization space, and (2) LLMs inherently encode rich optimization knowledge and can reason about architectural characteristics, which can effectively guide this decision process. Concretely, we first model tensor computation optimization as a Markov Decision Process (MDP), enabling unconstrained graph transformations over pre-defined scheduling rules. To efficiently explore the vast transformation space, we introduce an architecture-aware LLM-guided Monte Carlo Tree Search (MCTS). The LLM shapes the prior probability distribution over candidate transformations, guiding the search direction toward promising program sketches and parameter configurations. To adapt to concrete hardware and workloads, we propose an architecture-aware prior adaptation mechanism that distills natural-language heuristics from a lightweight offline stage. We conducted comprehensive experiments for representative subgraphs and LLMs on NVIDIA A100 and H100. Regarding subgraphs, QiMeng-Tensify averagely outperforms PyTorch, TensorRT, TVM, Triton, FlashAttention, Welder, Mirage, and Reasoning Compiler by , and , respectively. For LLM workloads, QiMeng-Tensify achieves average speedups of 1.56×, 1.22× and 1.30× over PyTorch, TensorRT-LLM and Mirage on the A100, and and on the H100, respectively. Results well demonstrate that QiMeng-Tensify provides a generalizable paradigm for optimizing large-scale tensor computation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic ApproachShouyang Dong, Jun Bi, Di Huang, Jiaming Guo et al.OSDI 2025 · 6 citations
- QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language ModelsQirui Zhou, Yuanbo Wen, Ruizhi Chen, Ke Gao et al.AAAI 2025 · 7 citations
- REASONING COMPILER: LLM-Guided Optimizations for Efficient Model ServingAnnabelle Sujun Tang, Christopher Priebe, Rohan Mahapatra, Lianhui Qin et al.NeurIPS 2025 · 7 citations
- ATiM: Autotuning Tensor Programs for Processing-in-DRAMYongwon Shin, Dookyung Kang, Hyojin SungISCA 2025 · 2 citations
- EGG: An Expert-Guided Agent Framework for Kernel GenerationYaochen Han, Ke Fan, Hongxu Jiang, Wanqi Xu et al.ICML 2026
