Landmark-Guided Policy Optimization for Multi-Objective Language Model Selection
Marcio Monteiro, Weichen Li, Puyu Wang, Marius Kloft, Sophie Fellenz
Abstract
Selecting a pretrained large language model (LLM) to fine-tune for a task-specific dataset can be time-consuming and costly. With several candidate models available to choose from, varying in size, architecture, and pretraining data, finding the best model for a specific task often involves extensive trial and error. In addition, the "best" model may not necessarily be the one with the lowest test loss, as practical considerations such as deployment costs, inference throughput, and limited search budgets might also play crucial roles. To address this, we introduce LAMPS (LAnguage Model Pareto Selection), a novel and open-source multi-objective AutoML framework that meta-learns a resource allocation policy to efficiently identify (or approximate) the Pareto front of candidate LLMs for a task-specific dataset. It is based on two key ideas: (1) landmark fine-tuning, which generates early performance indicators of the candidate models, and (2) meta-learning via reinforcement learning, which learns an effective selection policy from historical performance data (a meta-dataset). Our results show that, on held-out datasets, LAMPS reduces search time by an average of 73% compared to exhaustive search, while still covering more than 99% of the optimal target space hypervolume.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d24ce4c8-7590-4c15-9b71-0649130d3e72Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodBiao Zhang, Zhongtao Liu, Colin Cherry, Orhan FiratICLR 2024 · 271 citations
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 241 citations
- LoRA vs Full Fine-tuning: An Illusion of EquivalenceReece Shuttleworth, Jacob Andreas, Antonio Torralba, Pratyusha SharmaNeurIPS 2025 · 152 citations
Related papers
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoMLErnesto Luis Estevanell-Valladares, Suilan Estevez-Velarde, Yoan Gutiérrez, Andrés Montoyo et al.EMNLP 2025
- Pareto Prompt OptimizationGuang Zhao, Byung-Jun Yoon, Gilchan Park, Shantenu Jha et al.ICLR 2025
- Hardware Generation with High Flexibility using Reinforcement Learning Enhanced LLMsYifang Zhao, Weimin Fu, Shijie Li, Yi-Xiang Hu et al.DAC 2025 · 1 citation
- Optimization Methods for Personalizing Large Language Models through Retrieval AugmentationAlireza Salemi, Surya Kallumadi, Hamed ZamaniSIGIR 2024 · 52 citations
- AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision TasksZekang Yang, Wang Zeng, Sheng Jin, Chen Qian et al.AAAI 2025 · 18 citations
