Inverse Contextual Bandits: Learning How Behavior Evolves over Time
Alihan Hüyük, Daniel Jarrett, Mihaela van der Schaar
Abstract
Understanding a decision-maker's priorities by observing their behavior is critical for transparency and accountability in decision processes, such as in healthcare. Though conventional approaches to policy learning almost invariably assume stationarity in behavior, this is hardly true in practice: Medical practice is constantly evolving as clinical professionals fine-tune their knowledge over time. For instance, as the medical community's understanding of organ transplantations has progressed over the years, a pertinent question is: How have actual organ allocation policies been evolving? To give an answer, we desire a policy learning method that provides interpretable representations of decision-making, in particular capturing an agent's non-stationary knowledge of the world, as well as operating in an offline manner. First, we model the evolving behavior of decision-makers in terms of contextual bandits, and formalize the problem of Inverse Contextual Bandits (ICB). Second, we propose two concrete algorithms as solutions, learning parametric and nonparametric representations of an agent's behavior. Finally, using both real and simulated data for liver transplantations, we illustrate the applicability and explainability of our method, as well as benchmarking and validating its accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee9106b8-07ca-48b7-994d-34e3fceb3fc1Cited by top-tier papers7
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 50 citations
- Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RLHao Sun, Alihan Hüyük, Mihaela van der SchaarICLR 2024 · 48 citations
- Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of ExamplesHao Sun, Alihan Hüyük, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2023 · 13 citations
- Inverse Online Learning: Understanding Non-Stationary and Reactionary PoliciesAlex J. Chan, Alicia Curth, Mihaela van der SchaarICLR 2022 · 8 citations
- Online Decision MediationDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarNeurIPS 2022 · 5 citations
Builds on16
- Online Bayesian Goal Inference for Boundedly Rational Planning AgentsTan Zhi-Xuan, Jordyn L. Mann, Tom Silver, Josh Tenenbaum et al.NeurIPS 2020 · 122 citations
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 113 citations
- Strictly Batch Imitation Learning by Energy-based Distribution MatchingDaniel Jarrett, Ioana Bica, Mihaela van der SchaarNeurIPS 2020 · 74 citations
- OrganITE: Optimal transplant donor organ offering using an individual treatment effectJeroen Berrevoets, James Jordon, Ioana Bica, Alexander Gimson et al.NeurIPS 2020 · 51 citations
- What Did You Think Would Happen? Explaining Agent Behaviour through Intended OutcomesHerman Yau, Chris Russell, Simon HadfieldNeurIPS 2020 · 44 citations
Related papers
- Explaining by Imitating: Understanding Decisions by Interpretable Policy LearningAlihan Hüyük, Daniel Jarrett, Cem Tekin, Mihaela van der SchaarICLR 2021 · 22 citations
- Learning "What-if" Explanations for Sequential Decision-MakingIoana Bica, Daniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICLR 2021 · 5 citations
- Closing the loop in medical decision support by understanding clinical decision-making: A case study on organ transplantationYuchao Qin, Fergus Imrie, Alihan Hüyük, Daniel Jarrett et al.NeurIPS 2021 · 7 citations
- Inverse Decision Modeling: Learning Interpretable Representations of BehaviorDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICML 2021 · 30 citations
- PAC-Bayesian Offline Contextual Bandits With GuaranteesOtmane Sakhi, Pierre Alquier, Nicolas ChopinICML 2023 · 23 citations
