Multiple Choice Learning of Low-Rank Adapters for Language Modeling
Victor Letzelter, Hugo Malard, Mathieu Fontaine, Gaël Richard, Slim Essid, Andrei Bursuc, Patrick Perez
Abstract
We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference time. Traditional language modeling is an intrinsically ill-posed problem: given a context, multiple ``futures'' may be equally plausible. Our approach leverages Multiple Choice Learning (MCL) and the winner-takes-all loss to efficiently handle ambiguity through Low-Rank Adaptation. We provide a theoretical interpretation of applying MCL to language modeling, assuming the data is generated from a mixture of distributions. We illustrate the proposed approach using mixtures of Markov chains. We then demonstrate with experiments on audio and visual captioning, as well as machine translation, that our method achieves high diversity and relevance in generated outputs. We release the code for applying LoRA-MCL to a wide range of language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 207213b9-51fc-4ed8-a9ca-a152712833d3Cited by top-tier papers1
Ask how each one uses itBuilds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- Taming Sparsely Activated Transformer with Stochastic ExpertsSimiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim et al.ICLR 2022 · 144 citations
Related papers
- Winner-takes-all for Multivariate Probabilistic Time Series ForecastingAdrien Cortés, Rémi Rehm, Victor LetzelterICML 2025
- Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealingDavid Perera, Victor Letzelter, Théo Mariotte, Adrien Cortés et al.NeurIPS 2024 · 14 citations
- Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysisVictor Letzelter, Mathieu Fontaine, Mickaël Chen, Patrick Pérez et al.NeurIPS 2023 · 15 citations
- MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersJiajun Cao, Yuan Zhang, Tao Huang, Ming Lu et al.CVPR 2025
- JailbreakLoRA: Your Downloaded LoRA from Sharing Platforms might be UnsafeFanjunduo Wei, Zhenheng Tang, Rongfei Zeng, Tongliang Liu et al.ICLR 2026
