Batched Low-Rank Adaptation of Foundation Models
Yeming Wen, Swarat Chaudhuri
Abstract
Low-Rank Adaptation (LoRA) has recently gained attention for fine-tuning foundation models by incorporating trainable low-rank matrices, thereby reducing the number of trainable parameters. While LoRA offers numerous advantages, its applicability for real-time serving to a diverse and global user base is constrained by its incapability to handle multiple task-specific adapters efficiently. This imposes a performance bottleneck in scenarios requiring personalized, task-specific adaptations for each incoming request. To mitigate this constraint, we introduce Fast LoRA (FLoRA), a framework in which each input example in a minibatch can be associated with its unique low-rank adaptation weights, allowing for efficient batching of heterogeneous requests. We empirically demonstrate that FLoRA retains the performance merits of LoRA, showcasing competitive results on the MultiPL-E code generation benchmark spanning over 8 languages and a multilingual speech recognition task across 6 languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06a9184a-690c-477b-9a72-fd3f9360fbebCited by top-tier papers17
- SVFT: Parameter-Efficient Fine-Tuning with Singular VectorsVijay Lingam, Atula Neerkaje, Aditya Vavre, Aneesh Shetty et al.NeurIPS 2024 · 72 citations
- Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?Xi Chen, Kaituo Feng, Changsheng Li, Xunhao Lai et al.NeurIPS 2025 · 48 citations
- PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularizationYao Ni, Shan Zhang, Piotr KoniuszNeurIPS 2024 · 25 citations
- 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and ComposabilityBaohao Liao, Christof MonzNeurIPS 2024 · 14 citations
- Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt AdaptationAbhinav Jain, Swarat Chaudhuri, Thomas W. Reps, Christopher M. JermaineNeurIPS 2024 · 10 citations
Builds on27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- qa-FLoRA: Data-free query-adaptive Fusion of LoRAs for LLMsShreya Shukla, Aditya Sriram, Milinda Kuppur Narayanaswamy, Hiteshi JainAAAI 2026
- PLoRA: Efficient Concurrent LoRA Training for Large Language ModelsMinghao Yan, Zhuang Wang, Zhen Jia, Shivaram Venkataraman et al.ICML 2026 · 5 citations
- MELoRA: Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-TuningPengjie Ren, Chengshun Shi, Shiguang Wu, Mengqi Zhang et al.ACL 2024
- Compress then Serve: Serving Thousands of LoRA Adapters with Little OverheadRickard Brüel Gabrielsson, Jiacheng Zhu, Onkar Bhardwaj, Leshem Choshen et al.ICML 2025
- Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank AdaptationShiwei Li, Xiandi Luo, Haozhao Wang, Xing Tang et al.NeurIPS 2025 · 10 citations
