Lune

ICML2026顶会

Deep Reinforcement Learning Finds Bayes-Nash Equilibrium in Competitive Newsvendor Problems

Kassian Köck, Fabian Raoul Pieroth, Martin Bichler

出版方
2026年份

摘要

We investigate learning dynamics in competitive newsvendor games, a class of continuousaction games with strategic substitutes. Despite established equilibrium properties, convergence of independent learning algorithms in repeated general-sum play remains uncertain. We analyze structural properties under complete and incomplete information, deriving closed-form equilibria for a symmetric complete-information benchmark with perfect substitution. Our main theoretical contribution proves strict monotonicity in both complete-information and Bayesian models with private costs, ensuring equilibrium uniqueness. This provides convergence guarantees for variational-inequality-based algorithms. Numerical experiments using deep reinforcement learning agents with Proximal Policy Optimization empirically demonstrate convergence to Nash and Bayesian Nash equilibria, verified by equilibrium checks. These results establish a foundation for applying deep reinforcement learning in competitive inventory management.

We analyze both a complete-information setting, in which agents' cost parameters are fixed, and a Bayesian setting with private cost information. The complete-information model admits a closed-form symmetric Nash equilibrium and serves as a transparent baseline for studying learning dynamics. The Bayesian model captures environments in which agents face heterogeneous and privately known costs, and equilibrium strategies are functions over finite type spaces. These two settings correspond naturally to repeated interaction among fixed competitors and to environments with changing participants.

Existing work characterizes equilibria in both settings but typically relies on implicit or numerical representations. More importantly, prior analyses do not explain why independent learning algorithms should converge in these games. This paper makes three contributions.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext c9bebf04-a1f2-4c0e-a754-c0f43fb115a1

它引用的顶会 Paper3

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖