Preference-Optimized Retrieval and Ranking for Efficient Multimodal Recommendation
Zhenrui Yue, Huimin Zeng, Yueqi Wang, Julian J. McAuley, Dong Wang
Abstract
Large multimodal models (LMMs) exhibit enhanced capabilities in understanding and generating both textual and visual content. By leveraging item metadata, LMMs are also applied for recommendation and demonstrate improvements across diverse scenarios. However, the majority of existing methods explore static item attributes without considering additional contextual information (e.g., price, brand). Moreover, overlooking the interaction between the retrieval and ranking stages may lead to suboptimal solutions for fine-grained recommendations. In this work, we introduce PRIME: preference-optimized retrieval and ranking for efficient multimodal recommendation. PRIME operates in two stages: (i) a lightweight retriever identifies potential candidate items; (ii) an LMM learns to rank the retrieved candidates with detailed user history and multimodal features (e.g., text and image attributes). These features are incorporated into a carefully designed prompt, facilitating fine-grained transition patterns for user preference understanding. To optimize the inference efficiency of PRIME, we introduce verbalizer-based inference, which computes ranking scores for all candidate items in a single forward pass. Furthermore, we employ the LMM ranker to provide feedback on sampled candidate sets, enabling online preference optimization that refines the retriever model and improves the alignment between retrieval and ranking. As a result, PRIME can capture subtle user intentions and efficiently rank candidate items with minimal inference costs. Extensive experiments show the effectiveness and efficiency of PRIME, which consistently achieves superior performance over baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0ea0934-ec73-4343-9f00-b770d1008540Cited by top-tier papers1
Ask how each one uses itBuilds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan et al.SIGIR 2026 · 1 citation
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang et al.SIGIR 2025 · 11 citations
- Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationYuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang et al.AAAI 2025 · 68 citations
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential RecommendationYu Wang, Yonghui Yang, Le Wu, Yi Zhang et al.SIGIR 2026 · 9 citations
- PMG : Personalized Multimodal Generation with Large Language ModelsXiaoteng Shen, Rui Zhang, Xiaoyan Zhao, Jieming Zhu et al.WWW 2024 · 40 citations
