Characterizing Optimal Mixed Policies: Where to Intervene and What to Observe
Sanghack Lee, Elias Bareinboim
Abstract
© 2020 Neural information processing systems foundation. All rights reserved.Intelligent agents are continuously faced with the challenge of optimizing a policy based on what they can observe (see) and which actions they can take (do) in the environment where they are deployed. Most policies can be parametrized in terms of these two dimensions, i.e., as a function of what can be seen and done given a certain situation, which we call a mixed policy. In this paper, we investigate several properties of the class of mixed policies and provide an efficient and effective characterization, including optimality and non-redundancy. Specifically, we introduce a graphical criterion to identify unnecessary contexts for a set of actions, leading to a natural characterization of non-redundancy of mixed policies. We then derive sufficient conditions under which one strategy can dominate the other with respect to their maximum achievable expected rewards (optimality). This characterization leads to a fundamental understanding of the space of mixed policies and a possible refinement of the agents strategy so that it converges to the optimum faster and more robustly. One surprising result of the causal characterization is that the agent following a more standard approach—intervening on all intervenable variables and observing all available contexts—may be hurting itself, and will never achieve an optimal performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- The Causal-Neural Connection: Expressiveness, Learnability, and InferenceKevin Xia, Kai-Zhan Lee, Yoshua Bengio, Elias BareinboimNeurIPS 2021 · 158 citations
- Agent Incentives: A Causal PerspectiveTom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega et al.AAAI 2021 · 66 citations
- Adaptively Exploiting d-Separators with Causal BanditsBlair L. Bilodeau, Linbo Wang, Daniel M. RoyNeurIPS 2022 · 25 citations
- On Learning Necessary and Sufficient Causal GraphsHengrui Cai, Yixin Wang, Michael I. Jordan, Rui SongNeurIPS 2023 · 19 citations
- Combinatorial Causal BanditsShi Feng, Wei ChenAAAI 2023 · 16 citations
Builds on2
Related papers
- Reinforcement Learning in Reward-Mixing MDPsJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 23 citations
- Reinforcement Learning of Causal Variables Using Mediation AnalysisTue Herlau, Rasmus LarsenAAAI 2022 · 8 citations
- Towards Estimating Bounds on the Effect of Policies under Unobserved ConfoundingAlexis Bellot, Silvia ChiappaNeurIPS 2024 · 6 citations
- Structural Causal Bandits under Markov EquivalenceMin Woo Park, Andy Arditi, Elias Bareinboim, Sanghack LeeNeurIPS 2025 · 3 citations
- Balancing Context Length and Mixing Times for Reinforcement Learning at ScaleMatthew Riemer, Khimya Khetarpal, Janarthanan Rajendran, Sarath ChandarNeurIPS 2024 · 7 citations
