Learning Value Functions from Undirected State-only Experience
Matthew Chang, Arjun Gupta, Saurabh Gupta
Abstract
This paper tackles the problem of learning value functions from undirected stateonly experience (state transitions without action labels i.e. (s, s , r) tuples). We first theoretically characterize the applicability of Q-learning in this setting. We show that tabular Q-learning in discrete Markov decision processes (MDPs) learns the same value function under any arbitrary refinement of the action space. This theoretical result motivates the design of Latent Action Q-learning or LAQ, an offline RL method that can learn effective value functions from state-only experience. Latent Action Q-learning (LAQ) learns value functions using Q-learning on discrete latent actions obtained through a latent-variable future prediction model. We show that LAQ can recover value functions that have high correlation with value functions learned using ground truth actions. Value functions learned using LAQ lead to sample efficient acquisition of goal-directed behavior, can be used with domain-specific low-level controllers, and facilitate transfer across embodiments. Our experiments in 5 environments ranging from 2D grid world to 3D visual navigation in realistic environments demonstrate the benefits of LAQ over simpler alternatives, imitation learning oracles, and competing methods. * denotes equal contribution. Project website: https://matthewchang.github.io/latent action qlearning site/ . 1 We assume rt is observed. Reward can often be sparsely labeled in observation streams with low effort.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00fb25fe-e434-4530-81d9-97c5b9618191Cited by top-tier papers6
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- Reinforcement Learning from Passive Data via Latent IntentionsDibya Ghosh, Chethan Anand Bhateja, Sergey LevineICML 2023 · 69 citations
- Look Ma, No Hands! Agent-Environment Factorization of Egocentric VideosMatthew Chang, Aditya Prakash, Saurabh GuptaNeurIPS 2023 · 26 citations
- Learning Transferable Interaction Primitives from Game Videos for Humanoid LocomotionXiangming Zhu, Huayu Deng, Haoran Zhao, Yiwei Hao et al.ICML 2026
- Learning Video-Conditioned Policy on Unlabelled Data with Joint Embedding Predictive TransformerHao Luo, Zongqing LuICLR 2025
Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- Semantic Visual Navigation by Watching YouTube VideosMatthew Chang, Arjun Gupta, Saurabh GuptaNeurIPS 2020 · 108 citations
Related papers
- Latent Action Pretraining from VideosSeonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo et al.ICLR 2025
- Planning from Pixels using Inverse Dynamics ModelsKeiran Paster, Sheila A. McIlraith, Jimmy BaICLR 2021 · 44 citations
- Action-Free Offline-To-Online RL via Discretised State PoliciesNatinael Solomon Neggatu, Jeremie Houssineau, Giovanni MontanaICLR 2026
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- Latent Wasserstein Adversarial Imitation LearningSiqi Yang, Kai Yan, Alex Schwing, Yu-Xiong WangICLR 2026 · 1 citation
