Data-Driven Offline Decision-Making via Invariant Representation Learning
Han Qi, Yi Su, Aviral Kumar, Sergey Levine
Abstract
The goal in offline data-driven decision-making is synthesize decisions that optimize a black-box utility function, using a previously-collected static dataset, with no active interaction. These problems appear in many forms: offline reinforcement learning (RL), where we must produce actions that optimize the long-term reward, bandits from logged data, where the goal is to determine the correct arm, and offline model-based optimization (MBO) problems, where we must find the optimal design provided access to only a static dataset. A key challenge in all these settings is distributional shift: when we optimize with respect to the input into a model trained from offline data, it is easy to produce an out-of-distribution (OOD) input that appears erroneously good. In contrast to prior approaches that utilize pessimism or conservatism to tackle this problem, in this paper, we formulate offline data-driven decision-making as domain adaptation, where the goal is to make accurate predictions for the value of optimized decisions ("target domain"), when training only on the dataset ("source domain"). This perspective leads to invariant objective models (IOM), our approach for addressing distributional shift by enforcing invariance between the learned representations of the training dataset and optimized decisions. In IOM, if the optimized decisions are too different from the training dataset, the representation will be forced to lose much of the information that distinguishes good designs from bad ones, making all choices seem mediocre. Critically, when the optimizer is aware of this representational tradeoff, it should choose not to stray too far from the training distribution, leading to a natural trade-off between distributional shift and learning performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cee26e7a-985b-4d0c-85d9-652787e72326Cited by top-tier papers14
- Preference-Guided Diffusion for Multi-Objective Offline OptimizationYashas Annadani, Syrine Belakaria, Stefano Ermon, Stefan Bauer et al.NeurIPS 2025 · 12 citations
- ``Noisier'’ Noise Contrastive Estimation is (Almost) Maximum LikelihoodPeiyu Yu, Dinghuai Zhang, Hengzhi He, Xiaojian Ma et al.ICLR 2026 · 11 citations
- Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement LearningQi Wang, Junming Yang, Yunbo Wang, Xin Jin et al.NeurIPS 2024 · 10 citations
- Cliqueformer: Model-Based Optimization with Structured TransformersJakub Grudzien Kuba, Pieter Abbeel, Sergey LevineAAAI 2026 · 5 citations
- ROOT: Rethinking Offline Optimization as Distributional Translation via Probabilistic BridgeCuong Dao, The Hung Tran, Phi Le Nguyen, Truong Thao Nguyen et al.NeurIPS 2025 · 4 citations
Builds on18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Model-based reinforcement learning for biological sequence designChristof Angermüller, David Dohan, David Belanger, Ramya Deshpande et al.ICLR 2020 · 159 citations
Related papers
- Conservative Objective Models for Effective Offline Model-Based OptimizationBrandon Trabucco, Aviral Kumar, Xinyang Geng, Sergey LevineICML 2021 · 119 citations
- Representation Balancing Offline Model-based Reinforcement LearningByung-Jun Lee, Jongmin Lee, Kee-Eung KimICLR 2021 · 8 citations
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang et al.AAAI 2025 · 2 citations
- Confidence-Conditioned Value Functions for Offline Reinforcement LearningJoey Hong, Aviral Kumar, Sergey LevineICLR 2023 · 4 citations
- Offline RL Policies Should Be Trained to be AdaptiveDibya Ghosh, Anurag Ajay, Pulkit Agrawal, Sergey LevineICML 2022 · 62 citations
