Modeling Human Exploration Through Resource-Rational Reinforcement Learning
Marcel Binz, Eric Schulz
Abstract
Equipping artificial agents with useful exploration mechanisms remains a challenge to this day. Humans, on the other hand, seem to manage the trade-off between exploration and exploitation effortlessly. In the present article, we put forward the hypothesis that they accomplish this by making optimal use of limited computational resources. We study this hypothesis by meta-learning reinforcement learning algorithms that sacrifice performance for a shorter description length (defined as the number of bits required to implement the given algorithm). The emerging class of models captures human exploration behavior better than previously considered approaches, such as Boltzmann exploration, upper confidence bound algorithms, and Thompson sampling. We additionally demonstrate that changing the description length in our class of models produces the intended effects: reducing description length captures the behavior of brain-lesioned patients while increasing it mirrors cognitive development during adolescence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- In-Context Impersonation Reveals Large Language Models' Strengths and BiasesLeonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz et al.NeurIPS 2023 · 259 citations
- Meta-in-context learning in large language modelsJulian Coda-Forno, Marcel Binz, Zeynep Akata, Matt M. Botvinick et al.NeurIPS 2023 · 81 citations
- Reinforcement Learning with Simple Sequence PriorsTankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz et al.NeurIPS 2023 · 18 citations
- In-Context Learning Agents Are Asymmetric Belief UpdatersJohannes A. Schubert, Akshay K. Jagadish, Marcel Binz, Eric SchulzICML 2024 · 16 citations
- Predictive auxiliary objectives in deep RL mimic learning in the brainChing Fang, Kim StachenfeldICLR 2024 · 16 citations
Builds on2
Related papers
- Dynamic allocation of limited memory resources in reinforcement learningNisheet Patel, Luigi Acerbi, Alexandre PougetNeurIPS 2020 · 6 citations
- Efficient Exploration in Resource-Restricted Reinforcement LearningZhihai Wang, Taoxing Pan, Qi Zhou, Jie WangAAAI 2023 · 14 citations
- Variational Bayesian Reinforcement Learning with Regret BoundsBrendan O'DonoghueNeurIPS 2021 · 48 citations
- Using natural language and program abstractions to instill human inductive biases in machinesSreejan Kumar, Carlos G. Correa, Ishita Dasgupta, Raja Marjieh et al.NeurIPS 2022 · 34 citations
- Discovering Temporally-Aware Reinforcement Learning AlgorithmsMatthew Thomas Jackson, Chris Lu, Louis Kirsch, Robert Tjarko Lange et al.ICLR 2024 · 23 citations
