Online Decision Making with Generative Action Sets
Jianyu Xu, Vidhi Jain, Bryan Wilder, Aarti Singh
摘要
With advances in generative AI, decision-making agents can now dynamically create new actions during online learning, but action generation typically incurs costs that must be balanced against potential benefits. We study an online learning problem where an agent can generate new actions at any time step by paying a one-time cost, with these actions becoming permanently available for future use. The challenge lies in learning the optimal sequence of two-fold decisions: which action to take and when to generate new ones, further complicated by the triangular tradeoffs among exploitation, exploration and . To solve this problem, we propose a doubly-optimistic algorithm that employs Lower Confidence Bounds (LCB) for action selection and Upper Confidence Bounds (UCB) for action generation. Empirical evaluation on healthcare question-answering datasets demonstrates that our approach achieves favorable generation-quality tradeoffs compared to baseline strategies. From theoretical perspectives, we prove that our algorithm achieves the optimal regret of , providing the first sublinear regret bound for online learning with expanding action spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Growing Action SpacesGregory Farquhar, Laura Gustafson, Zeming Lin, Shimon Whiteson 等ICML 2020 · 被引用 48 次
- Non-stationary Bandits with KnapsacksShang Liu, Jiashuo Jiang, Xiaocheng LiNeurIPS 2022 · 被引用 34 次
- Exploratory Machine Learning with Unknown UnknownsPeng Zhao, Yu-Jie Zhang, Zhi-Hua ZhouAAAI 2021 · 被引用 29 次
- Almost Tight Bounds for Online Facility Location in the Random-Order ModelHaim Kaplan, David Naori, Danny RazSODA 2023 · 被引用 8 次
相关 Paper
- Efficient Optimistic Exploration in Linear-Quadratic Regulators via Lagrangian RelaxationMarc Abeille, Alessandro LazaricICML 2020 · 被引用 31 次
- Dynamical Linear BanditsMarco Mussi, Alberto Maria Metelli, Marcello RestelliICML 2023 · 被引用 3 次
- Upper Confidence Primal-Dual Reinforcement Learning for CMDP with Adversarial LossShuang Qiu, Xiaohan Wei, Zhuoran Yang, Jieping Ye 等NeurIPS 2020 · 被引用 65 次
- Dynamic Assortment Selection and Pricing with Censored Preference FeedbackJung-hun Kim, Min-hwan OhICLR 2025
- Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning ApproachJunzhe ZhangICML 2020 · 被引用 78 次
