Bayesian Reparameterization of Reward-Conditioned Reinforcement Learning with Energy-based Models
Wenhao Ding, Tong Che, Ding Zhao, Marco Pavone
Abstract
Recently, reward-conditioned reinforcement learning (RCRL) has gained popularity due to its simplicity, flexibility, and off-policy nature. However, we will show that current RCRL approaches are fundamentally limited and fail to address two critical challenges of RCRL -- improving generalization on high reward-to-go (RTG) inputs, and avoiding out-of-distribution (OOD) RTG queries during testing time. To address these challenges when training vanilla RCRL architectures, we propose Bayesian Reparameterized RCRL (BR-RCRL), a novel set of inductive biases for RCRL inspired by Bayes' theorem. BR-RCRL removes a core obstacle preventing vanilla RCRL from generalizing on high RTG inputs -- a tendency that the model treats different RTG inputs as independent values, which we term ``RTG Independence". BR-RCRL also allows us to design an accompanying adaptive inference method, which maximizes total returns while avoiding OOD queries that yield unpredictable behaviors in vanilla RCRL methods. We show that BR-RCRL achieves state-of-the-art performance on the Gym-Mujoco and Atari offline RL benchmarks, improving upon vanilla RCRL by up to 11%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5760d43b-60ed-4bd5-b39f-601c3876c587Cited by top-tier papers2
- OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement LearningYihang Yao, Zhepeng Cen, Wenhao Ding, Haohong Lin et al.NeurIPS 2024 · 16 citations
- A Tractable Inference Perspective of Offline RLXuejie Liu, Anji Liu, Guy Van den Broeck, Yitao LiangNeurIPS 2024 · 3 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
Related papers
- Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement LearningLong Ma, Fangwei Zhong, Yizhou WangICML 2025
- What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?Rui Yang, Lin Yong, Xiaoteng Ma, Hao Hu et al.ICML 2023 · 35 citations
- Entropy Regularized Task Representation Learning for Offline Meta-Reinforcement LearningMohammadreza Nakhaeinezhadfard, Aidan Scannell, Joni PajarinenAAAI 2025
- Offline Meta Reinforcement Learning with In-Distribution Online AdaptationJianhao Wang, Jin Zhang, Haozhe Jiang, Junyu Zhang et al.ICML 2023 · 16 citations
- Revisiting OOD Generalization in Programmatic RLAmirhossein Rajabpour, Kiarash Aghakasiri, Sandra Zilles, Levi LelisICML 2026
