Optimistic Task Inference for Behavior Foundation Models
Thomas Rupf, Marco Bagatella, Marin Vlastelica, Andreas Krause
Abstract
Behavior Foundation Models (BFMs) are capable of retrieving high-performing policy for any reward function specified directly at test-time, commonly referred to as zero-shot reinforcement learning (RL). While this is a very efficient process in terms of compute, it can be less so in terms of data: as a standard assumption, BFMs require computing rewards over a non-negligible inference dataset, assuming either access to a functional form of rewards, or significant labeling efforts. To alleviate these limitations, we tackle the problem of task inference purely through interaction with the environment at test-time. We propose OpTI-BFM, an optimistic decision criterion that directly models uncertainty over reward functions and guides BFMs in data collection for task inference. Formally, we provide a regret bound for well-trained BFMs through a direct connection to upper-confidence algorithms for linear bandits. Empirically, we evaluate OpTI-BFM on established zero-shot benchmarks, and observe that it enables successor-features-based BFMs to identify and optimize an unseen reward function in a handful of episodes with minimal compute overhead. Code is available at https://github.com/ThomasRupf/opti-bfm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f33d25f-40d4-4525-ae61-01f3271fa9e9Builds on9
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 140 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- Misspecified Gaussian Process Bandit OptimizationIlija Bogunovic, Andreas KrauseNeurIPS 2021 · 69 citations
- Information Directed Reward Learning for Reinforcement LearningDavid Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek et al.NeurIPS 2021 · 27 citations
Related papers
- Zero-Shot Adaptation of Behavioral Foundation Models to Unseen DynamicsMaksim Bobrin, Ilya Zisman, Alexander Nikulin, Vladislav Kurenkov et al.ICLR 2026 · 9 citations
- Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation ModelsPranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang et al.ICLR 2026 · 6 citations
- Proto Successor Measure: Representing the Behavior Space of an RL AgentSiddhant Agarwal, Harshit Sikchi, Peter Stone, Amy ZhangICML 2025
- Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation ModelsAndrea Tirinzoni, Ahmed Touati, Jesse Farebrother, Mateusz Guzek et al.ICLR 2025
- Improving Zero-Shot Offline RL via Behavioral Task SamplingNazim Bendib, Nicolas Perrin-Gilbert, Olivier SigaudICML 2026
