Token-Efficient Item Representation via Images for LLM Recommender Systems
Kibum Kim, Sein Kim, Hongseok Kang, Jiwan Kim, Heewoong Noh, Yeonjun In, Kanghoon Yoon, Jinoh Oh, Julian J. McAuley, Chanyoung Park
Abstract
Large Language Models (LLMs) have recently emerged as a powerful backbone for recommender systems. Existing LLM-based recommender systems take two different approaches for representing items in natural language, i.e., Attributebased Representation and Description-based Representation. In this work, we aim to address the trade-off between efficiency and effectiveness that these two approaches encounter, when representing items consumed by users. Based on our observation that there is a significant information overlap between images and descriptions associated with items, we propose a novel method, Image representation for LLM-based Recommender system (I-LLMRec). Our main idea is to leverage images as an alternative to lengthy textual descriptions for representing items, aiming at reducing token usage while preserving the rich semantic information of item descriptions. Through extensive experiments on real-world Amazon datasets, we demonstrate that I-LLMRec outperforms existing methods that leverage textual descriptions for representing items in both efficiency and effectiveness by leveraging images. Moreover, a further appeal of I-LLMRec is its ability to reduce sensitivity to noise in descriptions, leading to more robust recommendations. Our code is available at https://github.com/rlqja1107/torch-I-LLMRec . * Corresponding Author 1 While attributes are the high-level, general features in a few keywords (e.g., Apple), descriptions generally provide item-specific details (e.g., It is a slim metallic body, 13-inch Liquid Retina, and black keyboard...).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e30e340-57ac-45a1-9bb5-a6a12fee368fBuilds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- IDGenRec: LLM-RecSys Alignment with Textual ID LearningJuntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge et al.SIGIR 2024 · 45 citations
- Bridging Items and Language: A Transition Paradigm for Large Language Model-Based RecommendationXinyu Lin, Wenjie Wang, Yongqi Li, Fuli Feng et al.KDD 2024 · 27 citations
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim et al.KDD 2025 · 3 citations
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su et al.WWW 2024 · 385 citations
- Improving LLMs for Recommendation with Out-Of-Vocabulary TokensTing-Ji Huang, Jia-Qi Yang, Chunxu Shen, Kai-Qi Liu et al.ICML 2025
