Test-time Offline Reinforcement Learning on Goal-related Experience
Marco Bagatella, Mert Albaba, Jonas Hübotter, Georg Martius, Andreas Krause
Abstract
Foundation models compress a large amount of information in a single, large neural network, which can then be queried for individual tasks. There are strong parallels between this widespread framework and offline goal-conditioned reinforcement learning algorithms: a universal value function is trained on a large number of goals, and the policy is evaluated on a single goal in each test episode. Extensive research in foundation models has shown that performance can be substantially improved through test-time training, specializing the model to the current goal. We find similarly that test-time offline reinforcement learning on experience related to the test goal can lead to substantially better policies at modest compute costs. We propose a novel self-supervised data selection criterion, which selects transitions from an offline dataset according to their relevance to the current state and quality with respect to the evaluation goal. We demonstrate across a wide range of high-dimensional loco-navigation and manipulation tasks that fine-tuning a policy on the selected data for a few gradient steps leads to significant performance gains over standard offline pre-training. Our goal-conditioned test-time training (GC-TTT) algorithm applies this routine in a receding-horizon fashion during evaluation, adapting the policy to the current trajectory as it is being rolled out. Finally, we study compute allocation at inference, demonstrating that, at comparable costs, GC-TTT induces performance gains that are not achievable by scaling model size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe85e1d6-b817-406a-b6ed-d6e04b9b6a2fCited by top-tier papers3
- Learning to Discover at Test TimeMert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi et al.ICML 2026 · 73 citations
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation ModelsJonas Hübotter, Patrik Wolf, Aleksandr Shevchenko, Dennis Jüni et al.ICLR 2026 · 6 citations
- Dejavu: Towards Experience Feedback Learning for Embodied IntelligenceShaokai Wu, Yanbiao Ji, Qiuchang Li, Zhiyi Zhang et al.CVPR 2026 · 3 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- Flattening Hierarchies with Policy BootstrappingJohn L. Zhou, Jonathan C. KaoNeurIPS 2025 · 8 citations
- Test-Time Graph Search for Goal-Conditioned Reinforcement LearningEvgenii Opryshko, Junwei Quan, Claas Voelcker, Yilun Du et al.ICML 2026 · 6 citations
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- Self-Improving Embodied Foundation ModelsSeyed Kamyar Seyed Ghasemipour, Ayzaan Wahid, Jonathan Tompson, Pannag Sanketi et al.NeurIPS 2025 · 38 citations
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better InferenceYue Yu, Qiwei Di, Quanquan Gu, Dongruo ZhouICML 2026 · 3 citations
