Lune

ICML2026Top-tier venue

Representation Drift Compensation: A Near-Zero Inference Cost Enhancement for LLM Decomposition

Xinhao Huang, You-Liang Huang, Zeyi Wen

2026Year

Abstract

While low-rank decomposition offers potential for reducing LLM parameters, maintaining the original capabilities remains a significant challenge. In this work, we identify and formalize a key overlooked issue in LLM decomposition: representation drift. We show that approximation errors introduced by decomposition propagate and amplify non-linearly through the deep layers of the transformer architecture, progressively distorting internal representations and degrading downstream performance. To mitigate this, we introduce a conceptually simple but principled compensation mechanism, named ``Decomper'', that operates by suppressing error at its source. By learning to align the output distribution of decomposed transformer blocks with their original counterparts, our method effectively counteracts representation drift, achieving notable performance recovery with near-zero inference overhead. Extensive experiments in OPT, LLaMA-2/3, and Qwen exhibit remarkable improvements. For instance, on LLaMA-3-8B and OPT-13B at 40% compression, perplexity is reduced by more than 70% while reasoning task accuracy improves by over 10%. Our code is available at this URL.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f3f99ebc-1102-4696-a080-5d4486210612

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines