Inverse Rational Control with Partially Observable Continuous Nonlinear Dynamics
Minhae Kwon, Saurabh Daptardar, Paul R. Schrater, Xaq Pitkow
Abstract
A fundamental question in neuroscience is how the brain creates an internal model of the world to guide actions using sequences of ambiguous sensory information. This is naturally formulated as a reinforcement learning problem under partial observations, where an agent must estimate relevant latent variables in the world from its evidence, anticipate possible future states, and choose actions that optimize total expected reward. This problem can be solved by control theory, which allows us to find the optimal actions for a given system dynamics and objective function. However, animals often appear to behave suboptimally. Why? We hypothesize that animals have their own flawed internal model of the world, and choose actions with the highest expected subjective reward according to that flawed model. We describe this behavior as rational but not optimal. The problem of Inverse Rational Control (IRC) aims to identify which internal model would best explain an agent's actions. Our contribution here generalizes past work on Inverse Rational Control which solved this problem for discrete control in partially observable Markov decision processes. Here we accommodate continuous nonlinear dynamics and continuous actions, and impute sensory observations corrupted by unknown noise that is private to the animal. We first build an optimal Bayesian agent that learns an optimal policy generalized over the entire model space of dynamics and subjective rewards using deep reinforcement learning. Crucially, this allows us to compute a likelihood over models for experimentally observable action trajectories acquired from a suboptimal agent. We then find the model parameters that maximize the likelihood using gradient ascent. Our method successfully recovers the true model of rational agents. This approach provides a foundation for interpreting the behavioral and neural dynamics of animal brains during complex tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4824da1a-cdc7-4beb-b554-f6eadb33314dCited by top-tier papers12
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 50 citations
- Inverse Decision Modeling: Learning Interpretable Representations of BehaviorDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICML 2021 · 30 citations
- Explaining by Imitating: Understanding Decisions by Interpretable Policy LearningAlihan Hüyük, Daniel Jarrett, Cem Tekin, Mihaela van der SchaarICLR 2021 · 22 citations
- Amortized Inference with User SimulationsHee-Seung Moon, Antti Oulasvirta, Byungjoo LeeCHI 2023 · 20 citations
- Real-time 3D Target Inference via Biomechanical SimulationHee-Seung Moon, Yi-Chi Liao, Chenyu Li, Byungjoo Lee et al.CHI 2024 · 15 citations
Builds on4
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 162 citations
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 74 citations
- Meta-learning curiosity algorithmsFerran Alet, Martin F. Schneider, Tomás Lozano-Pérez, Leslie Pack KaelblingICLR 2020 · 67 citations
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 56 citations
Related papers
- Probabilistic inverse optimal control for non-linear partially observable systems disentangles perceptual uncertainty and behavioral costsDominik Straub, Matthias Schultheis, Heinz Koeppl, Constantin A. RothkopfNeurIPS 2023 · 8 citations
- Inverse Optimal Control Adapted to the Noise Characteristics of the Human Sensorimotor SystemMatthias Schultheis, Dominik Straub, Constantin A. RothkopfNeurIPS 2021 · 25 citations
- Stochastic Optimal Control and Estimation with Multiplicative and Internal NoiseFrancesco Damiani, Akiyuki Anzai, Jan Drugowitsch, Gregory C. DeAngelis et al.NeurIPS 2024 · 2 citations
- What do you know? Bayesian knowledge inference for navigating agentsMatthias Schultheis, Jana-Sophie Schönfeld, Constantin A. Rothkopf, Heinz KoepplNeurIPS 2025
- Inverse decision-making using neural amortized Bayesian actorsDominik Straub, Tobias F. Niehues, Jan Peters, Constantin A. RothkopfICLR 2025
