Adversarial Attacks on Online Learning to Rank with Click Feedback
Jinhang Zuo, Zhiyao Zhang, Zhiyong Wang, Shuai Li, Mohammad Hajiesmaili, Adam Wierman
摘要
Online learning to rank (OLTR) is a sequential decision-making problem where a learning agent selects an ordered list of items and receives feedback through user clicks. Although potential attacks against OLTR algorithms may cause serious losses in real-world applications, little is known about adversarial attacks on OLTR. This paper studies attack strategies against multiple variants of OLTR. Our first result provides an attack strategy against the UCB algorithm on classical stochastic bandits with binary feedback, which solves the key issues caused by bounded and discrete feedback that previous works can not handle. Building on this result, we design attack algorithms against UCB-based OLTR algorithms in position-based and cascade models. Finally, we propose a general attack strategy against any algorithm under the general click model. Each attack algorithm manipulates the learning agent into choosing the target attack item times, incurring a cumulative cost of . Experiments on synthetic and real data further validate the effectiveness of our proposed attack algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Online Corrupted User Detection and Regret MinimizationZhiyong Wang, Jize Xie, Tong Yu, Shuai Li 等NeurIPS 2023 · 被引用 8 次
- Online Learning to Rank under Corruption: A Robust Cascading Bandits ApproachFatemeh Ghaffari, Siddarth Sitaraman, Xutong Liu, Xuchuang Wang 等KDD 2026 · 被引用 1 次
- Stochastic Bandits Robust to Adversarial AttacksXuchuang Wang, Maoli Liu, Jinhang Zuo, Xutong Liu 等ICLR 2025
它引用的顶会 Paper5
- Adversarial Attacks on Adversarial BanditsYuzhe Ma, Zhijin ZhouICLR 2023 · 被引用 199 次
- Adversarial Attacks on Linear Contextual BanditsEvrard Garcelon, Baptiste Rozière, Laurent Meunier, Jean Tarbouriech 等NeurIPS 2020 · 被引用 60 次
- Minimax Regret for Cascading BanditsDaniel Vial, Sujay Sanghavi, Sanjay Shakkottai, R. SrikantNeurIPS 2022 · 被引用 18 次
- Observation-Free Attacks on Stochastic BanditsYinglun Xu, Bhuvesh Kumar, Jacob D. AbernethyNeurIPS 2021 · 被引用 13 次
- When Are Linear Stochastic Bandits Attackable?Huazheng Wang, Haifeng Xu, Hongning WangICML 2022 · 被引用 13 次
相关 Paper
- Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based ModelCheng Chen, Canzhe Zhao, Shuai LiAAAI 2022 · 被引用 5 次
- Unified Off-Policy Learning to Rank: a Reinforcement Learning PerspectiveZeyu Zhang, Yi Su, Hui Yuan, Yiran Wu 等NeurIPS 2023 · 被引用 9 次
- How do Online Learning to Rank Methods Adapt to Changes of Intent?Shengyao Zhuang, Guido ZucconSIGIR 2021 · 被引用 6 次
- UniRank: Unimodal Bandit Algorithms for Online RankingCamille-Sovanneary Gauthier, Romaric Gaudel, Élisa FromontICML 2022 · 被引用 6 次
- LT2R: Learning to Online Learning to Rank for Web SearchXiaokai Chu, Changying Hao, Shuaiqiang Wang, Dawei Yin 等ICDE 2024 · 被引用 1 次
