Online Reinforcement Learning for Mixed Policy Scopes
Junzhe Zhang, Elias Bareinboim
Abstract
Combination therapy refers to the use of multiple treatments – such as surgery, medication, and behavioral therapy - to cure a single disease, and has become a cornerstone for treating various conditions including cancer, HIV, and depression. All possible combinations of treatments lead to a collection of treatment regimens (i.e., policies) with mixed scopes, or what physicians could observe and which actions they should take depending on the context. In this paper, we investigate the online reinforcement learning setting for optimizing the policy space with mixed scopes. In particular, we develop novel online algorithms that achieve sublinear regret compared to an optimal agent deployed in the environment. The regret bound has a dependency on the maximal cardinality of the induced state-action space associated with mixed scopes. We further introduce a canonical representation for an arbitrary subset of interventional distributions given a causal diagram, which leads to a non-trivial, minimal representation of the model parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 930b2465-cfe6-4254-90fa-1ea71dd55b5eCited by top-tier papers3
- Structural Causal Bandits under Markov EquivalenceMin Woo Park, Andy Arditi, Elias Bareinboim, Sanghack LeeNeurIPS 2025 · 3 citations
- Contextual Causal Bayesian OptimisationVahan Arsenyan, Antoine Grosnit, Haitham Bou-Ammar, Arnak S. DalalyanICLR 2026 · 3 citations
- Counterfactual Structural Causal BanditsMin Woo Park, Sanghack LeeICLR 2026 · 1 citation
Builds on6
- The Causal-Neural Connection: Expressiveness, Learnability, and InferenceKevin Xia, Kai-Zhan Lee, Yoshua Bengio, Elias BareinboimNeurIPS 2021 · 158 citations
- Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning ApproachJunzhe ZhangICML 2020 · 78 citations
- Partial Counterfactual Identification from Observational and Experimental DataJunzhe Zhang, Jin Tian, Elias BareinboimICML 2022 · 77 citations
- Characterizing Optimal Mixed Policies: Where to Intervene and What to ObserveSanghack Lee, Elias BareinboimNeurIPS 2020 · 42 citations
- A Class of Algorithms for General Instrumental Variable ModelsNiki Kilbertus, Matt J. Kusner, Ricardo SilvaNeurIPS 2020 · 41 citations
Related papers
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 7 citations
- Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement LearningMirco Mutti, Riccardo De Santi, Marcello Restelli, Alexander Marx et al.ICLR 2024 · 6 citations
- Policy Optimization as Online Learning with Mediator FeedbackAlberto Maria Metelli, Matteo Papini, Pierluca D'Oro, Marcello RestelliAAAI 2021 · 11 citations
- Mutli-Armed Bandits with Network InterferenceAbhineet Agarwal, Anish Agarwal, Lorenzo Masoero, Justin WhitehouseNeurIPS 2024 · 4 citations
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez et al.NeurIPS 2022 · 63 citations
