The In-Sample Softmax for Offline Reinforcement Learning
Chenjun Xiao, Han Wang, Yangchen Pan, Adam White, Martha White
Abstract
Reinforcement learning (RL) agents can leverage batches of previously collected data to extract a reasonable control policy. An emerging issue in this offline RL setting, however, is that the bootstrapping update underlying many of our methods suffers from insufficient action-coverage: standard max operator may select a maximal action that has not been seen in the dataset. Bootstrapping from these inaccurate values can lead to overestimation and even divergence. There are a growing number of methods that attempt to approximate an in-sample max, that only uses actions well-covered by the dataset. We highlight a simple fact: it is more straightforward to approximate an in-sample softmax using only actions in the dataset. We show that policy iteration based on the in-sample softmax converges, and that for decreasing temperatures it approaches the in-sample max. We derive an In-Sample Actor-Critic (AC), using this in-sample softmax, and show that it is consistently better or comparable to existing offline RL methods, and is also wellsuited to fine-tuning. We release the code at github.com/hwang-ua/inac pytorch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef3a246b-4ffb-41c2-bafc-513d5bbbe00eCited by top-tier papers26
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
- Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion ModelYinan Zheng, Jianxiong Li, Dongjie Yu, Yujie Yang et al.ICLR 2024 · 72 citations
- ACT: Empowering Decision Transformer with Dynamic Programming via Advantage ConditioningChenxiao Gao, Chenyang Wu, Mingjun Cao, Rui Kong et al.AAAI 2024 · 31 citations
- Doubly Mild Generalization for Offline Reinforcement LearningYixiu Mao, Qi Wang, Yun Qu, Yuhang Jiang et al.NeurIPS 2024 · 30 citations
- Constrained Policy Optimization with Explicit Behavior Density For Offline Reinforcement LearningJing Zhang, Chi Zhang, Wenjia Wang, Bingyi JingNeurIPS 2023 · 19 citations
Builds on18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- ACTIVE: Offline Reinforcement Learning via Adaptive Imitation and In-sample V-EnsembleTianyuan Chen, Ronglong Cai, Faguo Wu, Xiao ZhangICLR 2025
- In-sample Actor Critic for Offline Reinforcement LearningHongchang Zhang, Yixiu Mao, Boyuan Wang, Shuncheng He et al.ICLR 2023
- Iteratively Refined Behavior Regularization for Offline Reinforcement LearningYi Ma, Jianye Hao, Xiaohan Hu, Yan Zheng et al.NeurIPS 2024 · 11 citations
- Uncertainty Weighted Actor-Critic for Offline Reinforcement LearningYue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind et al.ICML 2021 · 223 citations
- Harnessing Mixed Offline Reinforcement Learning Datasets via Trajectory WeightingZhang-Wei Hong, Pulkit Agrawal, Remi Tachet des Combes, Romain LarocheICLR 2023 · 1 citation
