GLaVis: Generative Latent Visual Queries for Multimodal Conversational Recommender Systems with MLLMs
Ting Yang, Li Chen, Yuhan Zhao
摘要
Multimodal Conversational Recommender Systems (MCRSs) hold the promise of capturing ineffable user preferences in diverse modalities and multi-turn interactions. However, existing paradigms based on Multimodal Large Language Models (MLLMs) are fundamentally constrained: They either are restricted by external retrievers or suffer from severe information bottlenecks by translating visual intents into intermediate textual descriptions. To transcend these limitations, we propose GLaVis (Generative Latent Visual Queries), an end-to-end generative visual retrieval framework, which enables an MLLM to directly produce continuous latent embeddings as visual queries for multimodal conversational recommendation. To bridge the semantic gap between MLLM representations and the visual space of the item, we integrate a lightweight visual query projector into the MLLM backbone. This module maps internal hidden states to the visual embedding space, using a reason-before-generation paradigm to ensure that the generated queries are deeply grounded in the multimodal context. Furthermore, to unify recommendation (i.e., visual retrieval) and natural language response generation within a single framework, we devise a multi-stage training strategy. By integrating meta-alignment via supervised fine-tuning with task-oriented reinforcement learning using Group Relative Policy Optimization (GRPO), we jointly optimize the dual objectives of precise visual alignment and response generation. Extensive experiments on two public datasets (MMD and FashionRec) demonstrate that GLaVis significantly outperforms state-of-the-art baselines in both visual recommendation accuracy and response quality. Our code is available at https://github.com/yt556677/GLaVis.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and AlgorithmsMiaosen Zhang, Yixuan Wei, Zhen Xing, Yifei Ma 等NeurIPS 2024 · 被引用 2 次
- Guiding Cross-Modal Representations with MLLM Priors via Preference AlignmentPengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu 等NeurIPS 2025 · 被引用 3 次
- Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationYuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang 等AAAI 2025 · 被引用 68 次
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang 等SIGIR 2025 · 被引用 11 次
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang 等ICLR 2026 · 被引用 80 次
