Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
Asaf B. Cassel, Haipeng Luo, Aviv Rosenberg, Dmitry Sotnikov
Abstract
In many real-world applications, it is hard to provide a reward signal in each step of a Reinforcement Learning (RL) process and more natural to give feedback when an episode ends. To this end, we study the recently proposed model of RL with Aggregate Bandit Feedback (RL-ABF), where the agent only observes the sum of rewards at the end of an episode instead of each reward individually. Prior work studied RL-ABF only in tabular settings, where the number of states is assumed to be small. In this paper, we extend ABF to linear function approximation and develop two efficient algorithms with near-optimal regret guarantees: a value-based optimistic algorithm built on a new randomization technique with a Q-functions ensemble, and a policy optimization algorithm that uses a novel hedging scheme over the ensemble.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Multi-turn Reinforcement Learning with Preference Human FeedbackLior Shani, Aviv Rosenberg, Asaf B. Cassel, Oran Lang et al.NeurIPS 2024 · 16 citations
- Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental LimitsFan Chen, Zeyu Jia, Alexander Rakhlin, Tengyang XieNeurIPS 2025 · 8 citations
- Warm-up Free Policy Optimization: Improved Regret in Linear Markov Decision ProcessesAsaf B. Cassel, Aviv RosenbergNeurIPS 2024 · 6 citations
- Primal-Dual Policy Optimization for Linear CMDPs with Adversarial LossesKihyun Yu, Seoungbin Bae, Dabeen LeeICLR 2026 · 2 citations
- Multi-objective Linear Reinforcement Learning with Lexicographic RewardsBo Xue, Dake Bu, Ji Cheng, Yuanyu Wan et al.ICML 2025
Builds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 304 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Learning Adversarial Markov Decision Processes with Bandit Feedback and Unknown TransitionChi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra et al.ICML 2020 · 117 citations
- Optimistic Policy Optimization with Bandit FeedbackLior Shani, Yonathan Efroni, Aviv Rosenberg, Shie MannorICML 2020 · 100 citations
Related papers
- Learning Adversarial Linear Mixture Markov Decision Processes with Bandit Feedback and Unknown TransitionCanzhe Zhao, Ruofeng Yang, Baoxiang Wang, Shuai LiICLR 2023
- Minimax Optimal Regret Bound for Reinforcement Learning with Trajectory FeedbackZihan Zhang, Yuxin Chen, Jason D. Lee, Simon Shaolei Du et al.ICML 2025
- Reinforcement Learning with Trajectory FeedbackYonathan Efroni, Nadav Merlis, Shie MannorAAAI 2021 · 48 citations
- Delay-Adapted Policy Optimization and Improved Regret for Adversarial MDP with Delayed Bandit FeedbackTal Lancewicki, Aviv Rosenberg, Dmitry SotnikovICML 2023 · 6 citations
- Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret AlgorithmSattar Vakili, Julia OlkhovskayaNeurIPS 2024 · 7 citations
