Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
Ekaterina Grishina, Stepan L. Kuznetsov, Askar Tsyganov, Ilya Ivanov, Daria Korovaitceva, Margarita Rusanova, Uliana Parkina, Alexander Derevyagin, Evgeny Frolov, Sergey Samsonov, Anton Lysenko
摘要
The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale. This drives a demand for a proper methodology for fair comparison between algorithms. Naive aggregation of performance metrics (e.g., averaging NDCG over benchmarks) can yield misleading rankings, undermining practical selection. To address this problem, we introduce a novel, data-driven ranking methodology based on Bradley-Terry (BT) model. We demonstrate that the obtained ranking depends on key dataset statistics. Additionally, we propose a novel metric for evaluating ranking consistency and demonstrate robustness of our ranking to incomplete data. Finally, we introduce a dataset-specific methodology for ranking algorithms on unseen datasets without running the models, relying on extensions of the Bradley–Terry framework, including BT trees and BT models with covariates.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Matrix-Free Two-to-Infinity and One-to-Two Norms EstimationAskar Tsyganov, Evgeny Frolov, Sergey Samsonov, Maxim RakhubaAAAI 2026 · 被引用 2 次
- Dobi-SVD: Differentiable SVD for LLM Compression and Some New PerspectivesQinsi Wang, Jinghan Ke, Masayoshi Tomizuka, Kurt Keutzer 等ICLR 2025
相关 Paper
- Preference-Based Dynamic Ranking Structure RecognitionNan Lu, Jian Shi, Xinyu TianNeurIPS 2025
- Rank Aggregation from Pairwise Comparisons in the Presence of Adversarial CorruptionsArpit Agarwal, Shivani Agarwal, Sanjeev Khanna, Prathamesh PatilICML 2020 · 被引用 11 次
- Rank Aggregation via Heterogeneous Thurstone Preference ModelsTao Jin, Pan Xu, Quanquan Gu, Farzad FarnoudAAAI 2020 · 被引用 19 次
- Rethinking Reward Modeling in Preference-based Large Language Model AlignmentHao Sun, Yunyi Shen, Jean-Francois TonICLR 2025
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
