Meta-trained agents implement Bayes-optimal agents
Vladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein, Miljan Martic, Shane Legg, Pedro A. Ortega
摘要
Memory-based meta-learning is a powerful technique to build agents that adapt fast to any task within a target distribution. A previous theoretical study has argued that this remarkable performance is because the meta-training protocol incentivises agents to behave Bayes-optimally. We empirically investigate this claim on a number of prediction and bandit tasks. Inspired by ideas from theoretical computer science, we show that meta-learned and Bayes-optimal agents not only behave alike, but they even share a similar computational structure, in the sense that one agent system can approximately simulate the other. Furthermore, we show that Bayes-optimal agents are fixed points of the meta-learning dynamics. Our results suggest that memory-based meta-learning might serve as a general technique for numerically approximating Bayes-optimal agents - that is, even for task distributions for which we currently don't possess tractable models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand 等ICML 2023 · 被引用 155 次
- Exploration in Approximate Hyper-State Space for Meta Reinforcement LearningLuisa M. Zintgraf, Leo Feng, Cong Lu, Maximilian Igl 等ICML 2021 · 被引用 45 次
- Cognitive Model Discovery via Disentangled RNNsKevin J. Miller, Maria K. Eckstein, Matt M. Botvinick, Zeb Kurth-NelsonNeurIPS 2023 · 被引用 39 次
- Learning Universal PredictorsJordi Grau-Moya, Tim Genewein, Marcus Hutter, Laurent Orseau 等ICML 2024 · 被引用 29 次
- Memory-Based Meta-Learning on Non-Stationary DistributionsTim Genewein, Grégoire Delétang, Anian Ruoss, Li Kevin Wenliang 等ICML 2023 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- Learning Not to Learn: Nature versus Nurture In SilicoRobert Tjarko Lange, Henning SprekelerAAAI 2022 · 被引用 10 次
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial ObservabilityPo-Chen Kuo, Han Hou, Will Dabney, Edgar Y. WalkerNeurIPS 2025
- Regularization Guarantees Generalization in Bayesian Reinforcement Learning through Algorithmic StabilityAviv Tamar, Daniel Soudry, Ev ZisselmanAAAI 2022 · 被引用 9 次
- Meta-Learning for Simple Regret MinimizationMohammad Javad Azizi, Branislav Kveton, Mohammad Ghavamzadeh, Sumeet KatariyaAAAI 2023 · 被引用 11 次
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke 等NeurIPS 2023 · 被引用 17 次
