Data-Driven Offline Decision-Making via Invariant Representation Learning
Han Qi, Yi Su, Aviral Kumar, Sergey Levine
摘要
The goal in offline data-driven decision-making is synthesize decisions that optimize a black-box utility function, using a previously-collected static dataset, with no active interaction. These problems appear in many forms: offline reinforcement learning (RL), where we must produce actions that optimize the long-term reward, bandits from logged data, where the goal is to determine the correct arm, and offline model-based optimization (MBO) problems, where we must find the optimal design provided access to only a static dataset. A key challenge in all these settings is distributional shift: when we optimize with respect to the input into a model trained from offline data, it is easy to produce an out-of-distribution (OOD) input that appears erroneously good. In contrast to prior approaches that utilize pessimism or conservatism to tackle this problem, in this paper, we formulate offline data-driven decision-making as domain adaptation, where the goal is to make accurate predictions for the value of optimized decisions ("target domain"), when training only on the dataset ("source domain"). This perspective leads to invariant objective models (IOM), our approach for addressing distributional shift by enforcing invariance between the learned representations of the training dataset and optimized decisions. In IOM, if the optimized decisions are too different from the training dataset, the representation will be forced to lose much of the information that distinguishes good designs from bad ones, making all choices seem mediocre. Critically, when the optimizer is aware of this representational tradeoff, it should choose not to stray too far from the training distribution, leading to a natural trade-off between distributional shift and learning performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Preference-Guided Diffusion for Multi-Objective Offline OptimizationYashas Annadani, Syrine Belakaria, Stefano Ermon, Stefan Bauer 等NeurIPS 2025 · 被引用 12 次
- ``Noisier'’ Noise Contrastive Estimation is (Almost) Maximum LikelihoodPeiyu Yu, Dinghuai Zhang, Hengzhi He, Xiaojian Ma 等ICLR 2026 · 被引用 11 次
- Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement LearningQi Wang, Junming Yang, Yunbo Wang, Xin Jin 等NeurIPS 2024 · 被引用 10 次
- Cliqueformer: Model-Based Optimization with Structured TransformersJakub Grudzien Kuba, Pieter Abbeel, Sergey LevineAAAI 2026 · 被引用 5 次
- ROOT: Rethinking Offline Optimization as Distributional Translation via Probabilistic BridgeCuong Dao, The Hung Tran, Phi Le Nguyen, Truong Thao Nguyen 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Model-based reinforcement learning for biological sequence designChristof Angermüller, David Dohan, David Belanger, Ramya Deshpande 等ICLR 2020 · 被引用 159 次
相关 Paper
- Conservative Objective Models for Effective Offline Model-Based OptimizationBrandon Trabucco, Aviral Kumar, Xinyang Geng, Sergey LevineICML 2021 · 被引用 119 次
- Representation Balancing Offline Model-based Reinforcement LearningByung-Jun Lee, Jongmin Lee, Kee-Eung KimICLR 2021 · 被引用 8 次
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang 等AAAI 2025 · 被引用 2 次
- Confidence-Conditioned Value Functions for Offline Reinforcement LearningJoey Hong, Aviral Kumar, Sergey LevineICLR 2023 · 被引用 4 次
- Offline RL Policies Should Be Trained to be AdaptiveDibya Ghosh, Anurag Ajay, Pulkit Agrawal, Sergey LevineICML 2022 · 被引用 62 次
