Online Reinforcement Learning for Mixed Policy Scopes
Junzhe Zhang, Elias Bareinboim
摘要
Combination therapy refers to the use of multiple treatments – such as surgery, medication, and behavioral therapy - to cure a single disease, and has become a cornerstone for treating various conditions including cancer, HIV, and depression. All possible combinations of treatments lead to a collection of treatment regimens (i.e., policies) with mixed scopes, or what physicians could observe and which actions they should take depending on the context. In this paper, we investigate the online reinforcement learning setting for optimizing the policy space with mixed scopes. In particular, we develop novel online algorithms that achieve sublinear regret compared to an optimal agent deployed in the environment. The regret bound has a dependency on the maximal cardinality of the induced state-action space associated with mixed scopes. We further introduce a canonical representation for an arbitrary subset of interventional distributions given a causal diagram, which leads to a non-trivial, minimal representation of the model parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Structural Causal Bandits under Markov EquivalenceMin Woo Park, Andy Arditi, Elias Bareinboim, Sanghack LeeNeurIPS 2025 · 被引用 3 次
- Contextual Causal Bayesian OptimisationVahan Arsenyan, Antoine Grosnit, Haitham Bou-Ammar, Arnak S. DalalyanICLR 2026 · 被引用 3 次
- Counterfactual Structural Causal BanditsMin Woo Park, Sanghack LeeICLR 2026 · 被引用 1 次
它引用的顶会 Paper6
- The Causal-Neural Connection: Expressiveness, Learnability, and InferenceKevin Xia, Kai-Zhan Lee, Yoshua Bengio, Elias BareinboimNeurIPS 2021 · 被引用 158 次
- Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning ApproachJunzhe ZhangICML 2020 · 被引用 78 次
- Partial Counterfactual Identification from Observational and Experimental DataJunzhe Zhang, Jin Tian, Elias BareinboimICML 2022 · 被引用 77 次
- Characterizing Optimal Mixed Policies: Where to Intervene and What to ObserveSanghack Lee, Elias BareinboimNeurIPS 2020 · 被引用 42 次
- A Class of Algorithms for General Instrumental Variable ModelsNiki Kilbertus, Matt J. Kusner, Ricardo SilvaNeurIPS 2020 · 被引用 41 次
相关 Paper
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 被引用 7 次
- Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement LearningMirco Mutti, Riccardo De Santi, Marcello Restelli, Alexander Marx 等ICLR 2024 · 被引用 6 次
- Policy Optimization as Online Learning with Mediator FeedbackAlberto Maria Metelli, Matteo Papini, Pierluca D'Oro, Marcello RestelliAAAI 2021 · 被引用 11 次
- Mutli-Armed Bandits with Network InterferenceAbhineet Agarwal, Anish Agarwal, Lorenzo Masoero, Justin WhitehouseNeurIPS 2024 · 被引用 4 次
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez 等NeurIPS 2022 · 被引用 63 次
