Landmark-Guided Policy Optimization for Multi-Objective Language Model Selection
Marcio Monteiro, Weichen Li, Puyu Wang, Marius Kloft, Sophie Fellenz
摘要
Selecting a pretrained large language model (LLM) to fine-tune for a task-specific dataset can be time-consuming and costly. With several candidate models available to choose from, varying in size, architecture, and pretraining data, finding the best model for a specific task often involves extensive trial and error. In addition, the "best" model may not necessarily be the one with the lowest test loss, as practical considerations such as deployment costs, inference throughput, and limited search budgets might also play crucial roles. To address this, we introduce LAMPS (LAnguage Model Pareto Selection), a novel and open-source multi-objective AutoML framework that meta-learns a resource allocation policy to efficiently identify (or approximate) the Pareto front of candidate LLMs for a task-specific dataset. It is based on two key ideas: (1) landmark fine-tuning, which generates early performance indicators of the candidate models, and (2) meta-learning via reinforcement learning, which learns an effective selection policy from historical performance data (a meta-dataset). Our results show that, on held-out datasets, LAMPS reduces search time by an average of 73% compared to exhaustive search, while still covering more than 99% of the optimal target space hypervolume.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodBiao Zhang, Zhongtao Liu, Colin Cherry, Orhan FiratICLR 2024 · 被引用 271 次
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 被引用 241 次
- LoRA vs Full Fine-tuning: An Illusion of EquivalenceReece Shuttleworth, Jacob Andreas, Antonio Torralba, Pratyusha SharmaNeurIPS 2025 · 被引用 152 次
相关 Paper
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoMLErnesto Luis Estevanell-Valladares, Suilan Estevez-Velarde, Yoan Gutiérrez, Andrés Montoyo 等EMNLP 2025
- Pareto Prompt OptimizationGuang Zhao, Byung-Jun Yoon, Gilchan Park, Shantenu Jha 等ICLR 2025
- Hardware Generation with High Flexibility using Reinforcement Learning Enhanced LLMsYifang Zhao, Weimin Fu, Shijie Li, Yi-Xiang Hu 等DAC 2025 · 被引用 1 次
- Optimization Methods for Personalizing Large Language Models through Retrieval AugmentationAlireza Salemi, Surya Kallumadi, Hamed ZamaniSIGIR 2024 · 被引用 52 次
- AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision TasksZekang Yang, Wang Zeng, Sheng Jin, Chen Qian 等AAAI 2025 · 被引用 18 次
