PerSim: Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents via Personalized Simulators
Anish Agarwal, Abdullah Omar Alomar, Varkey Alumootil, Devavrat Shah, Dennis Shen, Zhi Xu, Cindy Yang
Abstract
We consider offline reinforcement learning (RL) with heterogeneous agents under severe data scarcity, i.e., we only observe a single historical trajectory for every agent under an unknown, potentially sub-optimal policy. We find that the performance of state-of-the-art offline and model-based RL methods degrade significantly given such limited data availability, even for commonly perceived "solved" benchmark settings such as "MountainCar" and "CartPole". To address this challenge, we propose PerSim, a model-based offline RL approach which first learns a personalized simulator for each agent by collectively using the historical trajectories across all agents, prior to learning a policy. We do so by positing that the transition dynamics across agents can be represented as a latent function of latent factors associated with agents, states, and actions; subsequently, we theoretically establish that this function is well-approximated by a "low-rank" decomposition of separable agent, state, and action latent functions. This representation suggests a simple, regularized neural network architecture to effectively learn the transition dynamics per agent, even with scarce, offline data. We perform extensive experiments across several benchmark environments and RL methods. The consistent improvement of our approach, measured in terms of both state dynamics prediction and eventual reward, confirms the efficacy of our framework in leveraging limited historical data to simultaneously learn personalized policies across agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- LAPO: Latent-Variable Advantage-Weighted Policy Optimization for Offline Reinforcement LearningXi Chen, Ali Ghadirzadeh, Tianhe Yu, Jianhao Wang et al.NeurIPS 2022 · 52 citations
- CausalSim: A Causal Framework for Unbiased Trace-Driven SimulationAbdullah Omar Alomar, Pouya Hamadanian, Arash Nasr-Esfahany, Anish Agarwal et al.NSDI 2023 · 46 citations
- Counterfactual Identifiability of Bijective Causal ModelsArash Nasr-Esfahany, Mohammad Alizadeh, Devavrat ShahICML 2023 · 42 citations
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingYecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani et al.ICLR 2023 · 35 citations
- Learning to Extrapolate: A Transductive ApproachAviv Netanyahu, Abhishek Gupta, Max Simchowitz, Kaiqing Zhang et al.ICLR 2023
Builds on8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- Behavior Estimation from Multi-Source Data for Offline Reinforcement LearningGuoxi Zhang, Hisashi KashimaAAAI 2023 · 2 citations
- Semi-Supervised Offline Reinforcement Learning with Action-Free TrajectoriesQinqing Zheng, Mikael Henaff, Brandon Amos, Aditya GroverICML 2023 · 29 citations
- Offline Multi-Agent Reinforcement Learning with Knowledge DistillationWei-Cheng Tseng, Tsun-Hsuan Johnson Wang, Yen-Chen Lin, Phillip IsolaNeurIPS 2022 · 62 citations
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li et al.NeurIPS 2022 · 81 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
