Meta-trained agents implement Bayes-optimal agents
Vladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein, Miljan Martic, Shane Legg, Pedro A. Ortega
Abstract
Memory-based meta-learning is a powerful technique to build agents that adapt fast to any task within a target distribution. A previous theoretical study has argued that this remarkable performance is because the meta-training protocol incentivises agents to behave Bayes-optimally. We empirically investigate this claim on a number of prediction and bandit tasks. Inspired by ideas from theoretical computer science, we show that meta-learned and Bayes-optimal agents not only behave alike, but they even share a similar computational structure, in the sense that one agent system can approximately simulate the other. Furthermore, we show that Bayes-optimal agents are fixed points of the meta-learning dynamics. Our results suggest that memory-based meta-learning might serve as a general technique for numerically approximating Bayes-optimal agents - that is, even for task distributions for which we currently don't possess tractable models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 527a703f-813a-466e-afbd-732e76dab792Cited by top-tier papers16
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand et al.ICML 2023 · 155 citations
- Exploration in Approximate Hyper-State Space for Meta Reinforcement LearningLuisa M. Zintgraf, Leo Feng, Cong Lu, Maximilian Igl et al.ICML 2021 · 45 citations
- Cognitive Model Discovery via Disentangled RNNsKevin J. Miller, Maria K. Eckstein, Matt M. Botvinick, Zeb Kurth-NelsonNeurIPS 2023 · 39 citations
- Learning Universal PredictorsJordi Grau-Moya, Tim Genewein, Marcus Hutter, Laurent Orseau et al.ICML 2024 · 29 citations
- Memory-Based Meta-Learning on Non-Stationary DistributionsTim Genewein, Grégoire Delétang, Anian Ruoss, Li Kevin Wenliang et al.ICML 2023 · 16 citations
Builds on1
Related papers
- Learning Not to Learn: Nature versus Nurture In SilicoRobert Tjarko Lange, Henning SprekelerAAAI 2022 · 10 citations
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial ObservabilityPo-Chen Kuo, Han Hou, Will Dabney, Edgar Y. WalkerNeurIPS 2025
- Regularization Guarantees Generalization in Bayesian Reinforcement Learning through Algorithmic StabilityAviv Tamar, Daniel Soudry, Ev ZisselmanAAAI 2022 · 9 citations
- Meta-Learning for Simple Regret MinimizationMohammad Javad Azizi, Branislav Kveton, Mohammad Ghavamzadeh, Sumeet KatariyaAAAI 2023 · 11 citations
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2023 · 17 citations
