Inverse Contextual Bandits: Learning How Behavior Evolves over Time
Alihan Hüyük, Daniel Jarrett, Mihaela van der Schaar
摘要
Understanding a decision-maker's priorities by observing their behavior is critical for transparency and accountability in decision processes, such as in healthcare. Though conventional approaches to policy learning almost invariably assume stationarity in behavior, this is hardly true in practice: Medical practice is constantly evolving as clinical professionals fine-tune their knowledge over time. For instance, as the medical community's understanding of organ transplantations has progressed over the years, a pertinent question is: How have actual organ allocation policies been evolving? To give an answer, we desire a policy learning method that provides interpretable representations of decision-making, in particular capturing an agent's non-stationary knowledge of the world, as well as operating in an offline manner. First, we model the evolving behavior of decision-makers in terms of contextual bandits, and formalize the problem of Inverse Contextual Bandits (ICB). Second, we propose two concrete algorithms as solutions, learning parametric and nonparametric representations of an agent's behavior. Finally, using both real and simulated data for liver transplantations, we illustrate the applicability and explainability of our method, as well as benchmarking and validating its accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 被引用 50 次
- Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RLHao Sun, Alihan Hüyük, Mihaela van der SchaarICLR 2024 · 被引用 48 次
- Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of ExamplesHao Sun, Alihan Hüyük, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2023 · 被引用 13 次
- Inverse Online Learning: Understanding Non-Stationary and Reactionary PoliciesAlex J. Chan, Alicia Curth, Mihaela van der SchaarICLR 2022 · 被引用 8 次
- Online Decision MediationDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarNeurIPS 2022 · 被引用 5 次
它引用的顶会 Paper16
- Online Bayesian Goal Inference for Boundedly Rational Planning AgentsTan Zhi-Xuan, Jordyn L. Mann, Tom Silver, Josh Tenenbaum 等NeurIPS 2020 · 被引用 122 次
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 被引用 113 次
- Strictly Batch Imitation Learning by Energy-based Distribution MatchingDaniel Jarrett, Ioana Bica, Mihaela van der SchaarNeurIPS 2020 · 被引用 74 次
- OrganITE: Optimal transplant donor organ offering using an individual treatment effectJeroen Berrevoets, James Jordon, Ioana Bica, Alexander Gimson 等NeurIPS 2020 · 被引用 51 次
- What Did You Think Would Happen? Explaining Agent Behaviour through Intended OutcomesHerman Yau, Chris Russell, Simon HadfieldNeurIPS 2020 · 被引用 44 次
相关 Paper
- Explaining by Imitating: Understanding Decisions by Interpretable Policy LearningAlihan Hüyük, Daniel Jarrett, Cem Tekin, Mihaela van der SchaarICLR 2021 · 被引用 22 次
- Learning "What-if" Explanations for Sequential Decision-MakingIoana Bica, Daniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICLR 2021 · 被引用 5 次
- Closing the loop in medical decision support by understanding clinical decision-making: A case study on organ transplantationYuchao Qin, Fergus Imrie, Alihan Hüyük, Daniel Jarrett 等NeurIPS 2021 · 被引用 7 次
- Inverse Decision Modeling: Learning Interpretable Representations of BehaviorDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICML 2021 · 被引用 30 次
- PAC-Bayesian Offline Contextual Bandits With GuaranteesOtmane Sakhi, Pierre Alquier, Nicolas ChopinICML 2023 · 被引用 23 次
