Koopman Q-learning: Offline Reinforcement Learning via Symmetries of Dynamics
Matthias Weissenbacher, Samarth Sinha, Animesh Garg, Yoshinobu Kawahara
Abstract
Offline reinforcement learning leverages large datasets to train policies without interactions with the environment. The learned policies may then be deployed in real-world settings where interactions are costly or dangerous. Current algorithms over-fit to the training dataset and as a consequence perform poorly when deployed to out-of-distribution generalizations of the environment. We aim to address these limitations by learning a Koopman latent representation which allows us to infer symmetries of the system's underlying dynamic. The latter is then utilized to extend the otherwise static offline dataset during training; this constitutes a novel data augmentation framework which reflects the system's dynamic and is thus to be interpreted as an exploration of the environments phase space. To obtain the symmetries we employ Koopman theory in which nonlinear dynamics are represented in terms of a linear operator acting on the space of measurement functions of the system and thus symmetries of the dynamics may be inferred directly. We provide novel theoretical results on the existence and nature of symmetries relevant for control systems such as reinforcement learning settings. Moreover, we empirically evaluate our method on several benchmark offline reinforcement learning tasks and datasets including D4RL, Metaworld and Robosuite and find that by using our framework we consistently improve the state-of-the-art of model-free Q-learning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4df1924-2f09-4ae1-b700-de1c8237fec5Cited by top-tier papers7
- Look Beneath the Surface: Exploiting Fundamental Symmetry for Sample-Efficient Offline RLPeng Cheng, Xianyuan Zhan, Zhi-Hao Wu, Wenjia Zhang et al.NeurIPS 2023 · 23 citations
- Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement LearningVittorio Giammarino, Ruiqi Ni, Ahmed H. QureshiNeurIPS 2025 · 15 citations
- Efficient Dynamics Modeling in Interactive Environments with Koopman TheoryArnab Kumar Mondal, Siba Smarak Panigrahi, Sai Rajeswar, Kaleem Siddiqi et al.ICLR 2024 · 12 citations
- Information Shapes Koopman RepresentationXiaoyuan Cheng, Wenxuan Yuan, Yiming Yang, Yuanzhao Zhang et al.ICLR 2026 · 4 citations
- DreamSAC: Learning Hamiltonian World Models via Symmetry ExplorationJinzhou Tang, Fan Feng, Minghao Fu, Wenjun Lin et al.CVPR 2026 · 1 citation
Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- MDP Homomorphic Networks: Group Symmetries in Reinforcement LearningElise van der Pol, Daniel E. Worrall, Herke van Hoof, Frans A. Oliehoek et al.NeurIPS 2020 · 203 citations
Related papers
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- Sample Efficient Offline RL via T-Symmetry Enforced Latent State-StitchingPeng Cheng, Zhihao Wu, Jianxiong Li, Ziteng He et al.ICLR 2026
- Offline Trajectory Optimization for Offline Reinforcement LearningZiqi Zhao, Zhaochun Ren, Liu Yang, Yunsen Liang et al.KDD 2025
- Double Check Your State Before Trusting It: Confidence-Aware Bidirectional Offline Model-Based ImaginationJiafei Lyu, Xiu Li, Zongqing LuNeurIPS 2022 · 35 citations
- Constrained Latent Action Policies for Model-Based Offline Reinforcement LearningMarvin Alles, Philip Becker-Ehmck, Patrick van der Smagt, Maximilian KarlNeurIPS 2024 · 5 citations
