VisualLens: Personalization through Task-Agnostic Visual History
Wang Bill Zhu, Deqing Fu, Kai Sun, Yi Lu, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Yue Liu, Anuj Kumar, Xin Dong
摘要
Existing recommendation systems either rely on user interaction logs, such as online shopping history for shopping recommendations, or focus on text signals. However, item-based histories are not always accessible, and are not generalizable for multimodal recommendation. We hypothesize that a user's visual history -- comprising images from daily life -- can offer rich, task-agnostic insights into their interests and preferences, and thus be leveraged for effective personalization. To this end, we propose VisualLens, a novel framework that leverages multimodal large language models (MLLMs) to enable personalization using task-agnostic visual history. VisualLens extracts, filters, and refines a spectrum user profile from the visual history to support personalized recommendation. We created two new benchmarks, Google-Review-V and Yelp-V, with task-agnostic visual histories, and show that VisualLens improves over state-of-the-art item-based multimodal recommendations by 5-10% on Hit@3, and outperforms GPT-4o by 2-5%. Further analysis shows that VisualLens is robust across varying history lengths and excels at adapting to both longer histories and unseen content categories.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu 等ACL 2020 · 被引用 454 次
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 被引用 256 次
- Towards Universal Sequence Representation Learning for Recommender SystemsYupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li 等KDD 2022 · 被引用 245 次
- Multi-Behavior Hypergraph-Enhanced Transformer for Sequential RecommendationYuhao Yang, Chao Huang, Lianghao Xia, Yuxuan Liang 等KDD 2022 · 被引用 165 次
相关 Paper
- Multimodal Query Suggestion with Multi-Agent Reinforcement Learning from Human FeedbackZheng Wang, Bingzheng Gan, Wei ShiWWW 2024 · 被引用 18 次
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang 等SIGIR 2025 · 被引用 11 次
- Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationYuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang 等AAAI 2025 · 被引用 68 次
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan 等SIGIR 2026 · 被引用 1 次
- Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware RefinementBeibei Zhang, Yanan Lu, Ruobing Xie, Zongyi Li 等ACM MM 2025
