Learning Value Functions from Undirected State-only Experience
Matthew Chang, Arjun Gupta, Saurabh Gupta
摘要
This paper tackles the problem of learning value functions from undirected stateonly experience (state transitions without action labels i.e. (s, s , r) tuples). We first theoretically characterize the applicability of Q-learning in this setting. We show that tabular Q-learning in discrete Markov decision processes (MDPs) learns the same value function under any arbitrary refinement of the action space. This theoretical result motivates the design of Latent Action Q-learning or LAQ, an offline RL method that can learn effective value functions from state-only experience. Latent Action Q-learning (LAQ) learns value functions using Q-learning on discrete latent actions obtained through a latent-variable future prediction model. We show that LAQ can recover value functions that have high correlation with value functions learned using ground truth actions. Value functions learned using LAQ lead to sample efficient acquisition of goal-directed behavior, can be used with domain-specific low-level controllers, and facilitate transfer across embodiments. Our experiments in 5 environments ranging from 2D grid world to 3D visual navigation in realistic environments demonstrate the benefits of LAQ over simpler alternatives, imitation learning oracles, and competing methods. * denotes equal contribution. Project website: https://matthewchang.github.io/latent action qlearning site/ . 1 We assume rt is observed. Reward can often be sparsely labeled in observation streams with low effort.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 被引用 173 次
- Reinforcement Learning from Passive Data via Latent IntentionsDibya Ghosh, Chethan Anand Bhateja, Sergey LevineICML 2023 · 被引用 69 次
- Look Ma, No Hands! Agent-Environment Factorization of Egocentric VideosMatthew Chang, Aditya Prakash, Saurabh GuptaNeurIPS 2023 · 被引用 26 次
- Learning Transferable Interaction Primitives from Game Videos for Humanoid LocomotionXiangming Zhu, Huayu Deng, Haoran Zhao, Yiwei Hao 等ICML 2026
- Learning Video-Conditioned Policy on Unlabelled Data with Joint Embedding Predictive TransformerHao Luo, Zongqing LuICLR 2025
它引用的顶会 Paper7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 被引用 177 次
- Semantic Visual Navigation by Watching YouTube VideosMatthew Chang, Arjun Gupta, Saurabh GuptaNeurIPS 2020 · 被引用 108 次
相关 Paper
- Latent Action Pretraining from VideosSeonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo 等ICLR 2025
- Planning from Pixels using Inverse Dynamics ModelsKeiran Paster, Sheila A. McIlraith, Jimmy BaICLR 2021 · 被引用 44 次
- Action-Free Offline-To-Online RL via Discretised State PoliciesNatinael Solomon Neggatu, Jeremie Houssineau, Giovanni MontanaICLR 2026
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas 等ICML 2026 · 被引用 38 次
- Latent Wasserstein Adversarial Imitation LearningSiqi Yang, Kai Yan, Alex Schwing, Yu-Xiong WangICLR 2026 · 被引用 1 次
