Dynamic Service Fee Pricing under Strategic Behavior: Actions as Instruments and Phase Transition
Rui Ai, David Simchi-Levi, Feng Zhu
Abstract
We study a dynamic pricing problem for third-party platform service fees under strategic, far-sighted customers. In each time period, the platform sets a service fee based on historical data, observes the resulting transaction quantities, and collects revenue. The platform also monitors equilibrium prices influenced by both demand and supply. The objective is to maximize total revenue over a time horizon T . Our problem incorporates three practical challenges: (a) initially, the platform lacks knowledge of the demand side beforehand, necessitating a balance between exploring (learning the demand curve) and exploiting (maximizing revenue) simultaneously; (b) since only equilibrium prices and quantities are observable, traditional Ordinary Least Squares (OLS) estimators would be biased and inconsistent; (c) buyers are rational and strategic, seeking to maximize their consumer surplus and potentially misrepresenting their preferences. To address these challenges, we propose novel algorithmic solutions. Our approach involves: (i) a carefully designed active randomness injection to balance exploration and exploitation effectively; (ii) using non-i.i.d. actions as instrumental variables (IV) to consistently estimate demand; (iii) a low-switching cost design that promotes nearly truthful buyer behavior. We show an expected regret bound of O( √ T ∧ σ -2 S ) and demonstrate its optimality, up to logarithmic factors, with respect to both the time horizon T and the randomness in supply σ S . Despite its simplicity, our model offers valuable insights into the use of actions as estimation instruments, the benefits of low-switching pricing policies in mitigating strategic buyer behavior, and the role of supply randomness in facilitating exploration which leads to a phase transition of policy performance. one another, they still lead to excellent estimate of the demand curve. In Theorems 3.1, 3.2 and 4.3, we show our algorithms' optimal regret bounds.
• We discover that the randomness in supply can effectively assist us in learning the demand curve. This counterintuitive fact reveals why, in Theorem 3.2, we can achieve a regret of O(1), but in the case where there is no noise in supply, in Theorem 3.1, we can only expect a regret of O( √ T ). We explain in Section 4 via lower bound results (Theorem 4.1 and 4.2) that these orders of magnitude differences are fundamental and unavoidable. Additionally, we detailed the phase transition points of regret bounds regarding supply randomness.
• We investigate robust pricing in the absence of prior knowledge of the buyer's discount rate.
Our AaI and AAaI algorithms don't require the input of a discount rate to initiate, but can instead universally motivate buyers whose time has value to nearly truthfully report the demand curve. Specifically, our algorithm is also applicable in scenarios where the discount rate varies. The robustness of our algorithm is benign both in theory and applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f2f0192-4480-4a99-8af0-e1c9a7f84d13Builds on7
- Provably Efficient Reinforcement Learning with Linear Function Approximation under Adaptivity ConstraintsTianhao Wang, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 169 citations
- Online Pricing with Offline Data: Phase Transition and Inverse Square LawJinzhi Bu, David Simchi-Levi, Yunzong XuICML 2020 · 40 citations
- On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy EvaluationXiaohong Chen, Zhengling QiICML 2022 · 36 citations
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov et al.NeurIPS 2023 · 31 citations
- Reserve Pricing in Repeated Second-Price Auctions with Strategic BiddersAlexey DrutsaICML 2020 · 17 citations
Related papers
- When Demands Evolve Larger and Noisier: Learning and Earning in a Growing EnvironmentFeng Zhu, Zeyu ZhengICML 2020 · 15 citations
- Revenue Maximization Under Sequential Price Competition Via The Estimation Of -Concave Demand FunctionsDaniele Bracale, Moulinath Banerjee, Yuekai Sun, Cong ShiICLR 2026 · 2 citations
- Contextual Dynamic Pricing with Unknown Noise: Explore-then-UCB Strategy and Improved RegretsYiyun Luo, Will Wei Sun, Yufeng LiuNeurIPS 2022 · 19 citations
- Learning to price with resource constraints: from full information to machine-learned pricesRuicheng Ao, Jiashuo Jiang, David Simchi-LeviNeurIPS 2025 · 4 citations
- Pricing with Contextual Elasticity and Heteroscedastic ValuationJianyu Xu, Yu-Xiang WangICML 2024 · 3 citations
