Deep Reinforcement Learning Finds Bayes-Nash Equilibrium in Competitive Newsvendor Problems
Kassian Köck, Fabian Raoul Pieroth, Martin Bichler
Abstract
We investigate learning dynamics in competitive newsvendor games, a class of continuousaction games with strategic substitutes. Despite established equilibrium properties, convergence of independent learning algorithms in repeated general-sum play remains uncertain. We analyze structural properties under complete and incomplete information, deriving closed-form equilibria for a symmetric complete-information benchmark with perfect substitution. Our main theoretical contribution proves strict monotonicity in both complete-information and Bayesian models with private costs, ensuring equilibrium uniqueness. This provides convergence guarantees for variational-inequality-based algorithms. Numerical experiments using deep reinforcement learning agents with Proximal Policy Optimization empirically demonstrate convergence to Nash and Bayesian Nash equilibria, verified by equilibrium checks. These results establish a foundation for applying deep reinforcement learning in competitive inventory management.
We analyze both a complete-information setting, in which agents' cost parameters are fixed, and a Bayesian setting with private cost information. The complete-information model admits a closed-form symmetric Nash equilibrium and serves as a transparent baseline for studying learning dynamics. The Bayesian model captures environments in which agents face heterogeneous and privately known costs, and equilibrium strategies are functions over finite type spaces. These two settings correspond naturally to repeated interaction among fixed competitors and to environments with changing participants.
Existing work characterizes equilibria in both settings but typically relies on implicit or numerical representations. More importantly, prior analyses do not explain why independent learning algorithms should converge in these games. This paper makes three contributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9bebf04-a1f2-4c0e-a754-c0f43fb115a1Builds on3
- Chaos, Extremism and Optimism: Volume Analysis of Learning in GamesYun Kuen Cheung, Georgios PiliourasNeurIPS 2020 · 42 citations
- Beyond Monotonicity: On the Convergence of Learning Algorithms in Standard Auction GamesMartin Bichler, Stephan B. Lunowa, Matthias Oberlechner, Fabian R. Pieroth et al.AAAI 2025 · 5 citations
- On the Uniqueness of Bayesian Coarse Correlated Equilibria in Standard First-Price and All-Pay AuctionsMete Seref Ahunbay, Martin BichlerSODA 2025 · 5 citations
Related papers
- Computing Optimal Equilibria and Mechanisms via Learning in Zero-Sum Extensive-Form GamesBrian Hu Zhang, Gabriele Farina, Ioannis Anagnostides, Federico Cacciamani et al.NeurIPS 2023 · 17 citations
- Convergence Analysis of No-Regret Bidding Algorithms in Repeated AuctionsZhe Feng, Guru Guruganesh, Christopher Liaw, Aranyak Mehta et al.AAAI 2021 · 31 citations
- Self-Play Q-Learners Can Provably Collude in the Iterated Prisoner's DilemmaQuentin Bertrand, Juan Agustin Duque, Emilio Calvano, Gauthier GidelICML 2025
- Multi-Agent Learning under Uncertainty: Recurrence vs. ConcentrationKyriakos Lotidis, Panayotis Mertikopoulos, Nicholas Bambos, José H. BlanchetNeurIPS 2025 · 1 citation
- Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded RationalityStefanos Leonardos, Georgios Piliouras, Kelly SpendloveNeurIPS 2021 · 43 citations
