Characterizing Optimal Mixed Policies: Where to Intervene and What to Observe
Sanghack Lee, Elias Bareinboim
摘要
© 2020 Neural information processing systems foundation. All rights reserved.Intelligent agents are continuously faced with the challenge of optimizing a policy based on what they can observe (see) and which actions they can take (do) in the environment where they are deployed. Most policies can be parametrized in terms of these two dimensions, i.e., as a function of what can be seen and done given a certain situation, which we call a mixed policy. In this paper, we investigate several properties of the class of mixed policies and provide an efficient and effective characterization, including optimality and non-redundancy. Specifically, we introduce a graphical criterion to identify unnecessary contexts for a set of actions, leading to a natural characterization of non-redundancy of mixed policies. We then derive sufficient conditions under which one strategy can dominate the other with respect to their maximum achievable expected rewards (optimality). This characterization leads to a fundamental understanding of the space of mixed policies and a possible refinement of the agents strategy so that it converges to the optimum faster and more robustly. One surprising result of the causal characterization is that the agent following a more standard approach—intervening on all intervenable variables and observing all available contexts—may be hurting itself, and will never achieve an optimal performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- The Causal-Neural Connection: Expressiveness, Learnability, and InferenceKevin Xia, Kai-Zhan Lee, Yoshua Bengio, Elias BareinboimNeurIPS 2021 · 被引用 158 次
- Agent Incentives: A Causal PerspectiveTom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega 等AAAI 2021 · 被引用 66 次
- Adaptively Exploiting d-Separators with Causal BanditsBlair L. Bilodeau, Linbo Wang, Daniel M. RoyNeurIPS 2022 · 被引用 25 次
- On Learning Necessary and Sufficient Causal GraphsHengrui Cai, Yixin Wang, Michael I. Jordan, Rui SongNeurIPS 2023 · 被引用 19 次
- Combinatorial Causal BanditsShi Feng, Wei ChenAAAI 2023 · 被引用 16 次
它引用的顶会 Paper2
相关 Paper
- Reinforcement Learning in Reward-Mixing MDPsJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 23 次
- Reinforcement Learning of Causal Variables Using Mediation AnalysisTue Herlau, Rasmus LarsenAAAI 2022 · 被引用 8 次
- Towards Estimating Bounds on the Effect of Policies under Unobserved ConfoundingAlexis Bellot, Silvia ChiappaNeurIPS 2024 · 被引用 6 次
- Structural Causal Bandits under Markov EquivalenceMin Woo Park, Andy Arditi, Elias Bareinboim, Sanghack LeeNeurIPS 2025 · 被引用 3 次
- Balancing Context Length and Mixing Times for Reinforcement Learning at ScaleMatthew Riemer, Khimya Khetarpal, Janarthanan Rajendran, Sarath ChandarNeurIPS 2024 · 被引用 7 次
