Joint Online Learning and Decision-making via Dual Mirror Descent
Alfonso Lobos, Paul Grigas, Zheng Wen
摘要
We consider an online revenue maximization problem over a finite time horizon subject to lower and upper bounds on cost. At each period, an agent receives a context vector sampled i.i.d. from an unknown distribution and needs to make a decision adaptively. The revenue and cost functions depend on the context vector as well as some fixed but possibly unknown parameter vector to be learned. We propose a novel offline benchmark and a new algorithm that mixes an online dual mirror descent scheme with a generic parameter learning process. When the parameter vector is known, we demonstrate an regret result as well an bound on the possible constraint violations. When the parameter is not known and must be learned, we demonstrate that the regret and constraint violations are the sums of the previous terms plus terms that directly depend on the convergence of the learning process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Online Learning under Budget and ROI Constraints via Weak AdaptivityMatteo Castiglioni, Andrea Celli, Christian KroerICML 2024 · 被引用 12 次
- Decoupling Learning and Decision-Making: Breaking the O(T) Barrier in Online Resource Allocation with First-Order MethodsWenzhi Gao, Chunlin Sun, Chenyu Xue, Yinyu YeICML 2024 · 被引用 3 次
- 3D-Learning: Diffusion-Augmented Distributionally Robust Decision-Focused LearningJiaqi Wen, Lei Fan, Jianyi YangINFOCOM 2026 · 被引用 1 次
- Nearly Optimal Competitive Ratio for Online Allocation Problems with Two-sided Resource Constraints and Finite RequestsQixin Zhang, Wenbing Ye, Zaiyi Chen, Haoyuan Hu 等ICML 2023 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- A Bandit Learning Algorithm and Applications to Auction DesignKim Thang NguyenNeurIPS 2020 · 被引用 4 次
- Gradient-Variation Bound for Online Convex Optimization with ConstraintsShuang Qiu, Xiaohan Wei, Mladen KolarAAAI 2023 · 被引用 6 次
- A Unifying Framework for Online Optimization with Long-Term ConstraintsMatteo Castiglioni, Andrea Celli, Alberto Marchesi, Giulia Romano 等NeurIPS 2022 · 被引用 59 次
- Online Learning with Knapsacks: the Best of Both WorldsMatteo Castiglioni, Andrea Celli, Christian KroerICML 2022 · 被引用 47 次
- Online Learning with Unknown ConstraintsKarthik Sridharan, Seung Won Wilson YooICML 2025
