Lune

ISCA2026顶会

QiMeng-Tensify: Scaling Up Tensor Computation Optimization via Architecture-Aware LLM-Guided MCTS

Shouyang Dong, Jun Bi, Yuanbo Wen, Xiyue Yu, Jianxing Xu, Guanglin Xu, Ling Li, Xuehai Zhou, Tianshi Chen, Qi Guo

2026年份

摘要

The growing scale and complexity of large language models (LLMs) have intensified the need for optimizing largescale tensor computations (e.g., self-attention and mixture-ofexperts) on hardware platforms. Existing solutions rely on either manual expert optimization or exploration-based autotuning methods. However, neither approach scales effectively for LLMs with hundreds, even thousands of operators and dynamic control flows, because of prohibitive optimization overheads or suboptimal performance. To address this problem, we present QiMeng-Tensify, the first framework that combines LLMs with sequential decision optimization for large-scale graph-level tensor computation. Our key insight is that: (1) tensor computation optimization can be formulated as a generalized sequential decision problem to enlarge the optimization space, and (2) LLMs inherently encode rich optimization knowledge and can reason about architectural characteristics, which can effectively guide this decision process. Concretely, we first model tensor computation optimization as a Markov Decision Process (MDP), enabling unconstrained graph transformations over pre-defined scheduling rules. To efficiently explore the vast transformation space, we introduce an architecture-aware LLM-guided Monte Carlo Tree Search (MCTS). The LLM shapes the prior probability distribution over candidate transformations, guiding the search direction toward promising program sketches and parameter configurations. To adapt to concrete hardware and workloads, we propose an architecture-aware prior adaptation mechanism that distills natural-language heuristics from a lightweight offline stage. We conducted comprehensive experiments for representative subgraphs and LLMs on NVIDIA A100 and H100. Regarding subgraphs, QiMeng-Tensify averagely outperforms PyTorch, TensorRT, TVM, Triton, FlashAttention, Welder, Mirage, and Reasoning Compiler by 6.49×,2.86×,1.68×,2.64×,1.27×,13.49×1.29×6.49 \times, 2.86 \times, 1.68 \times, 2.64 \times, 1.27 \times, 13.49 \times 1.29 \times, and 1.31×1.31 \times, respectively. For LLM workloads, QiMeng-Tensify achieves average speedups of 1.56×, 1.22× and 1.30× over PyTorch, TensorRT-LLM and Mirage on the A100, and 1.78×,1.29×1.78 \times, 1.29 \times and 1.30×1.30 \times on the H100, respectively. Results well demonstrate that QiMeng-Tensify provides a generalizable paradigm for optimizing large-scale tensor computation.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖