LLM-as-a-Judge for Reliable and Explainable Offline Evaluation in Top-K Recommendation
Yue Que, Junyi Zhou, Xiaokun Zhang, Haiming Jin, Qiao Xiang, Chen Ma
摘要
Recommendation evaluation plays a crucial role in guiding the refinement and deployment of recommender systems. Most existing trials rely on offline evaluation using Top-K metrics computed over holdout user behaviors. However, we identify two fundamental limitations that undermine their ability to deliver reliable and explainable evaluations. Regarding reliability, offline evaluation treats observed user feedback as a proxy of true preferences and enforces rigid ID matching between the proxy and recommendation. In practice, feedback collections are inherently shaped by incomplete and biased item exposure, leading to distorted and unreliable assessments. Regarding explainability, Top-K metrics only establish numerical scores without offering meaningful insights to support them, thereby reinforcing the black-box nature of offline evaluation. In this paper, we propose a reliable and explainable LLM-as-a-Judge framework for offline recommendation evaluation. To enhance reliability, we introduce a semantic proxy from user textual behaviors to represent their true preferences. This proxy allows for more flexible matching between preferences and recommendations in the semantic space, rather than depending on the holdout feedback. To ensure explainability, the LLM Judge adopts a reasoning-then-scoring process to generate relevance judgments along with explicit rationale. Finally, we aggregate the individual scores into global Top-K metrics to quantify overall recommendation quality, and provide justification for each preference hit or miss. Extensive experiments demonstrate that the LLM Judge achieves solid reliability, explainability, and robustness in evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen 等SIGIR 2022 · 被引用 658 次
- Revisiting Graph Based Collaborative Filtering: A Linear Residual Graph Convolutional Network ApproachLei Chen, Le Wu, Richang Hong, Kun Zhang 等AAAI 2020 · 被引用 634 次
- Large Language Models can Accurately Predict Searcher PreferencesPaul Thomas, Seth Spielman, Nick Craswell, Bhaskar MitraSIGIR 2024 · 被引用 153 次
- Graph Trend Filtering Networks for RecommendationWenqi Fan, Xiaorui Liu, Wei Jin, Xiangyu Zhao 等SIGIR 2022 · 被引用 114 次
相关 Paper
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal EvaluatorsJongwoo Ko, Sungnyun Kim, Sungwoo Cho, Se-Young YunNeurIPS 2025 · 被引用 6 次
- Review-driven Personalized Preference Reasoning with Large Language Models for RecommendationJieyong Kim, Hyunseo Kim, Hyunjin Cho, SeongKu Kang 等SIGIR 2025 · 被引用 13 次
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development QualityChunyang Li, Yilun Zheng, Xinting Huang, Tianqing Fang 等ICLR 2026 · 被引用 14 次
- RecExplainer: Aligning Large Language Models for Explaining Recommendation ModelsYuxuan Lei, Jianxun Lian, Jing Yao, Xu Huang 等KDD 2024 · 被引用 18 次
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen 等ICLR 2025
