Lune

ICML2026顶会

TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling

Hongyaoxing Gu, Xinzhe Chen, LIJUAN HU, Liu fangfang

2026年份

摘要

Mixture-of-Experts (MoE) models achieve remarkable performance by sparsely activating specialized experts, yet their massive parameters in experts pose significant challenges for deployment. While low-rank quantization offers a promising route to compress MoE models, existing methods still incur nonnegligible memory overhead and inference latency. To address these limitations, we propose TILEQ, a fine-tuning-free post-training quantization (PTQ) method that employs 2D-tiling structured lowrank quantization to share low-rank factors across both input and output dimensions of MoE experts. Furthermore, we introduce an efficient inference technique for TILEQ that fuses multiple low-rank expert computations into a singlepass operation, significantly improving hardware utilization. Experiments show that TILEQ cuts down additional memory usage up to 10× and reduces inference latency to ∼5% while preserving state-of-the-art accuracy. Our code is

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper25

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖