Mechanism Design for LLM Fine-tuning with Multiple Reward Models
Haoran Sun, Yurong Chen, Siwei Wang, Chu Xu, Wei Chen, Xiaotie Deng
Abstract
Fine-tuning large language models (LLMs) to aggregate multiple preferences has attracted considerable research attention. With aggregation algorithms advancing, a potential economic scenario arises where fine-tuning services are provided to agents with different preferences. In this context, agents may benefit from strategically misreporting their preferences, but this could harm the aggregation performance. This paper addresses such incentive issues by framing it as a mechanism design problem: an LLM provider determines the fine-tuning objective (training rule) and the pricing scheme (payment rule) for agents. We primarily focus on training rules that maximize social welfare subject to certain regularizations, referred to as SW-Max rules. First, we show that under most circumstances, truthful reporting is sub-optimal with simply a SW-Max rule, thereby highlighting the necessity of payments. Second, we extend the VCG payment to implement SW-Max rules in dominant-strategy incentive compatibility (DSIC). We characterize sufficient conditions for payment equivalence and derive the necessary conditions for a payment rule to implement a SW-Max rule in DSIC and other principles. Third, we demonstrate that our mechanism is approximately DSIC with perturbed input, showcasing its robustness against the inevitable errors in real-world applications. Experiments on real LLM training results further confirm the practical implications of our results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91de94f6-3ee3-4241-aaf6-c2da6961455cCited by top-tier papers9
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He et al.ICLR 2026 · 211 citations
- Ad Auctions for LLMs via Retrieval Augmented GenerationMohammadTaghi Hajiaghayi, Sébastien Lahaie, Keivan Rezaei, Suho ShinNeurIPS 2024 · 31 citations
- Strategyproof Reinforcement Learning from Human FeedbackThomas Kleine Buening, Jiarui Gan, Debmalya Mandal, Marta KwiatkowskaNeurIPS 2025 · 10 citations
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic FrameworkKihyun Kim, Jiawei Zhang, Asuman Ozdaglar, Pablo A. ParriloICLR 2026 · 5 citations
- Pay for The Second-Best Service: A Game-Theoretic Approach against Dishonest LLM ProvidersYuhan Cao, Yu Wang, Sitong Liu, Miao Li et al.WWW 2026 · 3 citations
Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar et al.ICML 2024 · 212 citations
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 208 citations
Related papers
- Truthful Aggregation of LLMs with an Application to Online AdvertisingErmis Soumalias, Michael Curry, Sven SeukenNeurIPS 2025 · 44 citations
- Fine-tuning language models to find agreement among humans with diverse preferencesMichiel A. Bakker, Martin J. Chadwick, Hannah Sheahan, Michael Henry Tessler et al.NeurIPS 2022 · 349 citations
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi et al.ICML 2024 · 145 citations
- Enhancing Affine Maximizer Auctions with Correlation-Aware PaymentHaoran Sun, Xia Xuanzhi, Xu Chu, Xiaotie DengICML 2026
- Is Your LLM Overcharging You? Tokenization, Transparency, and IncentivesAnder Artola Velasco, Stratis Tsirtsis, Nastaran Okati, Manuel Gomez-RodriguezICML 2026 · 16 citations
