Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
Orin Levy, Aviv Rosenberg, Alon Peled-Cohen, Yishay Mansour
Abstract
We introduce OPO-CMDP, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of where and denote the state and action spaces, the horizon length, the number of episodes, and the finite function classes used to approximate the losses and dynamics, respectively. This is the first regret bound with optimal dependence on and , directly improving the current state-of-the-art (Qian, Hu, and Simchi-Levi, 2024). These results demonstrate that optimistic policy optimization provides a natural, computationally superior and theoretically near-optimal path for solving CMDPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 725007e6-7304-4a6e-bcf0-5f0bf33d3ef7Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Optimistic Policy Optimization with Bandit FeedbackLior Shani, Yonathan Efroni, Aviv Rosenberg, Shie MannorICML 2020 · 100 citations
- Policy Optimization in Adversarial MDPs: Improved Exploration via Dilated BonusesHaipeng Luo, Chen-Yu Wei, Chung-Wei LeeNeurIPS 2021 · 59 citations
- Learning Adversarial Markov Decision Processes with Delayed FeedbackTal Lancewicki, Aviv Rosenberg, Yishay MansourAAAI 2022 · 40 citations
- Optimism in Face of a Context: Regret Guarantees for Stochastic Contextual MDPOrin Levy, Yishay MansourAAAI 2023 · 13 citations
Related papers
- Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit FeedbackTal Lancewicki, Yishay MansourICML 2025
- Optimal Regret for Policy Optimization in Contextual BanditsOrin Levy, Yishay MansourICML 2026 · 1 citation
- Efficient Rate Optimal Regret for Adversarial Contextual MDPs Using Online Function ApproximationOrin Levy, Alon Cohen, Asaf B. Cassel, Yishay MansourICML 2023 · 10 citations
- Eluder-based Regret for Stochastic Contextual MDPsOrin Levy, Asaf B. Cassel, Alon Cohen, Yishay MansourICML 2024 · 10 citations
- Reinforcement Learning with History Dependent Dynamic ContextsGuy Tennenholtz, Nadav Merlis, Lior Shani, Martin Mladenov et al.ICML 2023 · 13 citations
