Learning to Mitigate AI Collusion on Economic Platforms
Gianluca Brero, Eric Mibuari, Nicolas Lepore, David C. Parkes
摘要
Algorithmic pricing on online e-commerce platforms raises the concern of tacit collusion, where reinforcement learning algorithms learn to set collusive prices in a decentralized manner and through nothing more than profit feedback. This raises the question as to whether collusive pricing can be prevented through the design of suitable "buy boxes," i.e., through the design of the rules that govern the elements of e-commerce sites that promote particular products and prices to consumers. In this paper, we demonstrate that reinforcement learning (RL) can also be used by platforms to learn buy box rules that are effective in preventing collusion by RL sellers. For this, we adopt the methodology of Stackelberg POMDPs, and demonstrate success in learning robust rules that continue to provide high consumer welfare together with sellers employing different behavior models or having out-of-distribution costs for goods. * Author order is alphabetical. This research is funded in part by Defense Advanced Research Projects Agency under Cooperative Agreement HR00111920029. The content of the information does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. This is approved for public release; distribution is unlimited. The work of G. Brero was also supported by the SNSF (Swiss National Science Foundation) under Fellowship P2ZHP1 191253. We thank Emilio Calvano and Justin Johnson for their availability to answer questions about their work and for guidance in replicating some of their results. We also thank Alon Eden, Matthias Gerstgrasser, and Alexander MacKay for for helpful discussions and feedback.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 等NeurIPS 2024 · 被引用 140 次
- Oracles & Followers: Stackelberg Equilibria in Deep Multi-Agent Reinforcement LearningMatthias Gerstgrasser, David C. ParkesICML 2023 · 被引用 27 次
- Understanding Strategic Platform Entry and Seller Exploration: A Stackelberg ModelGarrett Seo, Xintong Wang, David C. ParkesWWW 2026
它引用的顶会 Paper9
- What is Local Optimality in Nonconvex-Nonconcave Minimax Optimization?Chi Jin, Praneeth Netrapalli, Michael I. JordanICML 2020 · 被引用 381 次
- Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical StudyTanner Fiez, Benjamin Chasnov, Lillian J. RatliffICML 2020 · 被引用 144 次
- PreferenceNet: Encoding Human Preferences in Auction Design with Deep LearningNeehar Peri, Michael J. Curry, Samuel Dooley, John DickersonNeurIPS 2021 · 被引用 46 次
- Why Should I Trust You, Bellman? The Bellman Error is a Poor Replacement for Value ErrorScott Fujimoto, David Meger, Doina Precup, Ofir Nachum 等ICML 2022 · 被引用 43 次
- Certifying Strategyproof Auction NetworksMichael J. Curry, Ping-Yeh Chiang, Tom Goldstein, John DickersonNeurIPS 2020 · 被引用 37 次
相关 Paper
- Platform Behavior under Market Shocks: A Simulation Framework and Reinforcement-Learning Based StudyXintong Wang, Gary Qiurui Ma, Alon Eden, Clara Li 等WWW 2023 · 被引用 15 次
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang 等AAAI 2021 · 被引用 131 次
- BCORLE(λ): An Offline Reinforcement Learning and Evaluation Framework for Coupons Allocation in E-commerce MarketYang Zhang, Bo Tang, Qingyu Yang, Dou An 等NeurIPS 2021 · 被引用 23 次
- Determinants and Effects of Buy Box Suppression on AmazonJeffrey L. Gleason, Shuo Zhang, Christo WilsonWWW 2026
- Price Stability and Improved Buyer Utility with Presentation Design: A Theoretical Study of the Amazon Buy BoxOphir Friedler, Hu Fu, Anna R. Karlin, Ariana TangWWW 2025 · 被引用 1 次
