Inferring learning rules from animal decision-making
Zoe Ashwood, Nicholas A. Roy, Ji Hyun Bak, Jonathan W. Pillow
Abstract
How do animals learn? This remains an elusive question in neuroscience. Whereas reinforcement learning often focuses on the design of algorithms that enable artificial agents to efficiently learn new tasks, here we develop a modeling framework to directly infer the empirical learning rules that animals use to acquire new behaviors. Our method efficiently infers the trial-to-trial changes in an animal's policy, and decomposes those changes into a learning component and a noise component. Specifically, this allows us to: (i) compare different learning rules and objective functions that an animal may be using to update its policy; (ii) estimate distinct learning rates for different parameters of an animal's policy; (iii) identify variations in learning across cohorts of animals; and (iv) uncover trial-to-trial changes that are not captured by normative learning rules. After validating our framework on simulated choice data, we applied our model to data from rats and mice learning perceptual decision-making tasks. We found that certain learning rules were far more capable of explaining trial-to-trial changes in an animal's policy. Whereas the average contribution of the conventional REINFORCE learning rule to the policy update for mice learning the International Brain Laboratory's task was just 30%, we found that adding baseline parameters allowed the learning rule to explain 92% of the animals' policy updates under our model. Intriguingly, the best-fitting learning rates and baseline values indicate that an animal's policy update, at each trial, does not occur in the direction that maximizes expected reward. Understanding how an animal transitions from chance-level to high-accuracy performance when learning a new task not only provides neuroscientists with insight into their animals, but also provides concrete examples of biological learning algorithms to the machine learning community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Reinforcement Learning with Non-Exponential DiscountingMatthias Schultheis, Constantin A. Rothkopf, Heinz KoepplNeurIPS 2022 · 19 citations
- Beyond accuracy: generalization properties of bio-plausible temporal credit assignment rulesYuhan Helena Liu, Arna Ghosh, Blake A. Richards, Eric Shea-Brown et al.NeurIPS 2022 · 10 citations
- Model Based Inference of Synaptic Plasticity RulesYash Mehta, Danil Tyulmankov, Adithya Rajagopalan, Glenn Turner et al.NeurIPS 2024 · 9 citations
- Probabilistic inverse optimal control for non-linear partially observable systems disentangles perceptual uncertainty and behavioral costsDominik Straub, Matthias Schultheis, Heinz Koeppl, Constantin A. RothkopfNeurIPS 2023 · 8 citations
- What do you know? Bayesian knowledge inference for navigating agentsMatthias Schultheis, Jana-Sophie Schönfeld, Constantin A. Rothkopf, Heinz KoepplNeurIPS 2025
Related papers
- Flexible inference for animal learning rules using neural networksYuhan Helena Liu, Victor Geadah, Jonathan W. PillowNeurIPS 2025
- Distinguishing Learning Rules with Brain Machine InterfacesJacob P. Portes, Christian Schmid, James M. MurrayNeurIPS 2022 · 12 citations
- Inverse Reinforcement Learning with Switching Rewards and History Dependency for Characterizing Animal BehaviorsJingyang Ke, Feiyang Wu, Jiyi Wang, Jeffrey Markowitz et al.ICML 2025
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 50 citations
- A contrastive rule for meta-learningNicolas Zucchet, Simon Schug, Johannes von Oswald, Dominic Zhao et al.NeurIPS 2022 · 22 citations
