Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders
Yupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He, An Yan, Xiusi Chen, Julian J. McAuley
Abstract
Feature engineering has long been central to recommender systems, yet effectively leveraging textual item features remains challenging. Recent advances in large language models (LLMs) have enabled their use as semantic encoders for recommendation, but their roles and behaviors in this setting are still not well understood. Prior studies often rely on general-purpose embedding benchmarks (e.g., MTEB) when selecting LLMs, overlooking the unique characteristics of recommendation tasks. To address this gap, we introduce BLaIR, a comprehensive benchmark for evaluating LLMs as semantic encoders in recommendation scenarios. We contribute (1) a new large-scale Amazon Reviews 2023 dataset with over 570 million reviews and 48 million items, (2) a unified benchmark covering sequential recommendation, collaborative filtering, and product search, and (3) a new complex-query product search task featuring both semi-synthetic and real-world evaluation datasets. Experiments with 11 leading LLMs show that their rankings on BLaIR show little correlation with MTEB, highlighting the unique challenges of semantic encoding in recommendation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31959d46-9ec4-4f66-b482-311c911488a0Cited by top-tier papers36
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran et al.UIST 2024 · 143 citations
- Intent Representation Learning with Large Language Model for RecommendationYu Wang, Lei Sang, Yi Zhang, Yiwen ZhangSIGIR 2025 · 17 citations
- Reasoning over Semantic IDs Enhances Generative RecommendationYingzhi He, Yan Sun, Junfei Tan, Yuxin Chen et al.KDD 2026 · 15 citations
- Interactive Recommendation Agent with Active User CommandsJiakai Tang, Wen Chen, Yujie Luo, Xunke Xi et al.KDD 2026 · 14 citations
- Efficient LLM Serving for Agentic Workflows: A Data Systems PerspectiveNoppanat Wadlom, Junyi Shen, Yao LuSIGMOD 2026 · 14 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su et al.WWW 2024 · 385 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
Related papers
- LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential RecommendationYingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma et al.KDD 2025 · 2 citations
- LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial ApplicationJian Jia, Yipei Wang, Yan Li, Honggang Chen et al.AAAI 2025 · 37 citations
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling ViewJingzhe Liu, Liam Collins, Jiliang Tang, Tong Zhao et al.KDD 2026 · 17 citations
- Beyond Utility: Evaluating LLM as RecommenderChumeng Jiang, Jiayin Wang, Weizhi Ma, Charles L. A. Clarke et al.WWW 2025 · 22 citations
- SEAR: LLM-Powered Sequential Recommendation via Fusion of Collaborative, Semantic, and Rating InformationWei Guan, Jian Cao, Qiqi Cai, Jianqi Gao et al.WWW 2026
