L3Ms - Lagrange Large Language Models
Guneet S. Dhillon, Xingjian Shi, Yee Whye Teh, Alex Smola
Abstract
Supervised fine-tuning (SFT) and alignment of large language models (LLMs) are key steps in providing a good user experience. However, the concept of an appropriate alignment is inherently application-dependent, and current methods often rely on heuristic choices to drive optimization. In this work, we formulate SFT and alignment as a constrained optimization problem: the LLM is fine-tuned on a task while being required to meet application-specific requirements, without resorting to heuristics. To solve this, we propose Lagrange Large Language Models (L3Ms), which employ logarithmic barriers to enforce the constraints. This approach allows for the customization of L3Ms across diverse applications while avoiding heuristic-driven processes. We experimentally demonstrate the versatility and efficacy of L3Ms in achieving tailored alignments for various applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de256c4c-36aa-47f0-8b2d-559ddf67473fCited by top-tier papers1
Ask how each one uses itRelated papers
- Preference-Oriented Supervised Fine-Tuning: Favoring Target Model over Aligned Large Language ModelsYuchen Fan, Yuzhong Hong, Qiushi Wang, Junwei Bao et al.AAAI 2025 · 7 citations
- Data Selection for LLM Alignment Using Fine-Grained PreferencesJia Zhang, Yao Liu, Chen-Xi Zhang, Yi Liu et al.ICLR 2026 · 1 citation
- Alignment of Large Language Models with Constrained LearningBotong Zhang, Shuo Li, Ignacio Hounie, Osbert Bastani et al.NeurIPS 2025 · 12 citations
- Jailbreak Open-Sourced Large Language Models via Enforced DecodingHangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao et al.ACL 2024
- Beyond Imitation: Leveraging Fine-grained Quality Signals for AlignmentGeyang Guo, Ranchi Zhao, Tianyi Tang, Xin Zhao et al.ICLR 2024 · 44 citations
