GLaVis: Generative Latent Visual Queries for Multimodal Conversational Recommender Systems with MLLMs
Ting Yang, Li Chen, Yuhan Zhao
Abstract
Multimodal Conversational Recommender Systems (MCRSs) hold the promise of capturing ineffable user preferences in diverse modalities and multi-turn interactions. However, existing paradigms based on Multimodal Large Language Models (MLLMs) are fundamentally constrained: They either are restricted by external retrievers or suffer from severe information bottlenecks by translating visual intents into intermediate textual descriptions. To transcend these limitations, we propose GLaVis (Generative Latent Visual Queries), an end-to-end generative visual retrieval framework, which enables an MLLM to directly produce continuous latent embeddings as visual queries for multimodal conversational recommendation. To bridge the semantic gap between MLLM representations and the visual space of the item, we integrate a lightweight visual query projector into the MLLM backbone. This module maps internal hidden states to the visual embedding space, using a reason-before-generation paradigm to ensure that the generated queries are deeply grounded in the multimodal context. Furthermore, to unify recommendation (i.e., visual retrieval) and natural language response generation within a single framework, we devise a multi-stage training strategy. By integrating meta-alignment via supervised fine-tuning with task-oriented reinforcement learning using Group Relative Policy Optimization (GRPO), we jointly optimize the dual objectives of precise visual alignment and response generation. Extensive experiments on two public datasets (MMD and FashionRec) demonstrate that GLaVis significantly outperforms state-of-the-art baselines in both visual recommendation accuracy and response quality. Our code is available at https://github.com/yt556677/GLaVis.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and AlgorithmsMiaosen Zhang, Yixuan Wei, Zhen Xing, Yifei Ma et al.NeurIPS 2024 · 2 citations
- Guiding Cross-Modal Representations with MLLM Priors via Preference AlignmentPengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu et al.NeurIPS 2025 · 3 citations
- Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationYuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang et al.AAAI 2025 · 68 citations
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang et al.SIGIR 2025 · 11 citations
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang et al.ICLR 2026 · 80 citations
