Token-Efficient Item Representation via Images for LLM Recommender Systems
Kibum Kim, Sein Kim, Hongseok Kang, Jiwan Kim, Heewoong Noh, Yeonjun In, Kanghoon Yoon, Jinoh Oh, Julian J. McAuley, Chanyoung Park
摘要
Large Language Models (LLMs) have recently emerged as a powerful backbone for recommender systems. Existing LLM-based recommender systems take two different approaches for representing items in natural language, i.e., Attributebased Representation and Description-based Representation. In this work, we aim to address the trade-off between efficiency and effectiveness that these two approaches encounter, when representing items consumed by users. Based on our observation that there is a significant information overlap between images and descriptions associated with items, we propose a novel method, Image representation for LLM-based Recommender system (I-LLMRec). Our main idea is to leverage images as an alternative to lengthy textual descriptions for representing items, aiming at reducing token usage while preserving the rich semantic information of item descriptions. Through extensive experiments on real-world Amazon datasets, we demonstrate that I-LLMRec outperforms existing methods that leverage textual descriptions for representing items in both efficiency and effectiveness by leveraging images. Moreover, a further appeal of I-LLMRec is its ability to reduce sensitivity to noise in descriptions, leading to more robust recommendations. Our code is available at https://github.com/rlqja1107/torch-I-LLMRec . * Corresponding Author 1 While attributes are the high-level, general features in a few keywords (e.g., Apple), descriptions generally provide item-specific details (e.g., It is a slim metallic body, 13-inch Liquid Retina, and black keyboard...).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- IDGenRec: LLM-RecSys Alignment with Textual ID LearningJuntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge 等SIGIR 2024 · 被引用 45 次
- Bridging Items and Language: A Transition Paradigm for Large Language Model-Based RecommendationXinyu Lin, Wenjie Wang, Yongqi Li, Fuli Feng 等KDD 2024 · 被引用 27 次
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim 等KDD 2025 · 被引用 3 次
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su 等WWW 2024 · 被引用 385 次
- Improving LLMs for Recommendation with Out-Of-Vocabulary TokensTing-Ji Huang, Jia-Qi Yang, Chunxu Shen, Kai-Qi Liu 等ICML 2025
