Lune

ICLR2026顶会

CAR-LoRA: Training Compression-Aware and Robust LoRA Adapters for Evolving LLMs

Rana Muhammad Shahroz Khan, Zhen Tan, Ruichen Zhang, Hua Wei, Tianlong Chen, Charles Fleming

出版方
2026年份

摘要

The deployment of large language models (LLMs) for specialized tasks on resource-constrained edge devices like smartphones and sensors presents a significant scalability problem. To run on such hardware, these massive models must be compressed using techniques like quantization or pruning to reduce their memory and computational footprint. Concurrently, foundational LLMs are periodically updated by their developers with new data, making their internal parameters shift over time\textit{internal parameters shift over time}. While parameter-efficient methods like Low-Rank Adaptation (LoRA) streamline personalization by fine-tuning only a small fraction of parameters, the resulting adapters are brittle\textbf{brittle}; a LoRA trained for one specific compression scheme is incompatible with another, and an adapter trained on an older base model performs poorly on an updated one. This forces a costly cycle of retraining for each unique device and every new model release. To address this, we introduce a novel framework that creates a single, universally portable adapter that is both (i) compression-aware and (ii) temporally robust\textbf{\textit{(i)} compression-aware and \textit{(ii)} temporally robust}. We achieve this by augmenting the training process with a variety of simulated compression techniques during a single run, utilizing a quantized forward pass to build resilience while maintaining a full-precision backward pass for stable gradient optimization. This method yields a unified adapter robust to diverse compression artifacts and the subtle parameter shifts from model evolution\textit{This method yields a unified adapter robust to diverse compression artifacts and the subtle parameter shifts from model evolution}. Extensive experiments on models such as Llama-2, Llama-3.1, Gemma-2\texttt{Llama-2, Llama-3.1, Gemma-2}, and Mistral\texttt{Mistral} across reasoning benchmarks like SQA, MATH, and GSM8K\textit{SQA, MATH, and GSM8K} demonstrate that our single adapter achieves performance comparable to specialized adapters (e.g.\textit{e.g.}, QLoRA) that are individually retrained for each compression scheme. Furthermore, we show this single adapter maintains its high performance when applied to future, evolved versions of the base model, eliminating the need for periodic retraining. Our work pioneers an efficient paradigm for edge AI, creating portable model patches that bridge the gap between cloud-based personalization, the diverse hardware ecosystem, and the lifecycle of evolving LLMs.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖