Adaptive Mix Preference Optimization for Generative Recommendation
Junbo Qi, Yanyan Zou, Xuanhua Yang, Sulong Xu, Ying Sun, Shengjie Li
Abstract
Recommender systems aim to leverage user interaction signals to recommend items that users are likely to be interested in. Motivated by the success of Large Language Model (LLMs), generative recommendation (GR) has recently gained increasing attention, typically following a two-stage paradigm: supervised fine-tuning followed by preference alignment. However, aligning generative recommenders with users' personalized preferences remains challenging, as user feedback is inherently heterogeneous and uncertain. Different types of user interaction signals reflect varying levels of intent and should therefore be modeled differently. In this work, we propose Adaptive Mix Preference Optimization (AMPO), an adaptive alignment framework that mixes likelihood and preference objectives with self-calibrated, sample-wise confidence adjustment. AMPO introduces an adaptive target margin that leverages the model's own probability ratio to modulate optimization strength: confident pairs receive full margins that reinforce correct rankings, while uncertain pairs receive reduced margins that prevent overfitting to ambiguous signals. Additionally, AMPO incorporates negative log-likelihood regularization on preferred items to counteract likelihood displacement, a phenomenon where contrastive objectives cause preferred and non-preferred probabilities to collapse simultaneously. Such a design eliminates the need for a reference model, yielding up to 2.5x speedup and 35% memory reduction. Extensive experiments on public benchmarks and a large-scale industrial dataset demonstrate consistent improvements in ranking metrics. Online A/B tests on a major e-commerce platform further confirm statistically significant gains in click-through and conversion rates. The code is available at https://github.com/jumbo-q/ampo.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6accae4f-6d16-41c8-9767-9f72eae48291Related papers
- Align³GR: Unified Multi-Level Alignment for LLM-based Generative RecommendationWencai Ye, Mingjie Sun, Shuhang Chen, Wenjin Wu et al.AAAI 2026 · 2 citations
- What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal ContextZhongyu Ouyang, Qianlong Wen, Chunhui Zhang, Yanfang Ye et al.ACL 2026
- RMO: Towards Better LLM Alignment via Reshaping Reward Margin DistributionsYanchi Ru, Yue Huang, Xiangliang ZhangAAAI 2026 · 1 citation
- On Softmax Direct Preference Optimization for RecommendationYuxin Chen, Junfei Tan, An Zhang, Zhengyi Yang et al.NeurIPS 2024 · 126 citations
- MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference AlignmentTianze Wang, Dongnan Gui, Yifan Hu, Shuhang Lin et al.ICML 2025
