Transportability for Bandits with Data from Different Environments
Alexis Bellot, Alan Malek, Silvia Chiappa
Abstract
A unifying theme in the design of intelligent agents is to efficiently optimize a policy based on what prior knowledge of the problem is available and what actions can be taken to learn more about it. Bandits are a canonical instance of this task that has been intensely studied in the literature. Most methods, however, typically rely solely on an agent’s experimentation in a single environment (or multiple closely related environments). In this paper, we relax this assumption and consider the design of bandit algorithms from a combination of batch data and qualitative assumptions about the relatedness across different environments, represented in the form of causal models. In particular, we show that it is possible to exploit invariances across environments, wherever they may occur in the underlying causal model, to consistently improve learning. The resulting bandit algorithm has a sub-linear regret bound with an explicit dependency on a term that captures how informative related environments are for the task at hand; and may have substantially lower regret than experimentation-only bandit instances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02f4001b-a4e8-466a-bdb8-a87cb8a08ee5Cited by top-tier papers5
- Efficient Policy Evaluation Across Multiple Different Experimental DatasetsYonghan Jung, Alexis BellotNeurIPS 2024 · 4 citations
- Sequential Decision Making with Expert Demonstrations under Unobserved HeterogeneityVahid Balazadeh Meresht, Keertana Chidambaram, Viet Nguyen, Rahul G. Krishnan et al.NeurIPS 2024 · 3 citations
- Counterfactual Structural Causal BanditsMin Woo Park, Sanghack LeeICLR 2026 · 1 citation
- The Limits of Predicting Agents from BehaviourAlexis Bellot, Jonathan Richens, Tom EverittICML 2025
- Counterfactual Bootstrap for Robust Meta-Reinforcement LearningAi Bo, Junzhe Zhang, M. Cenk GursoyICML 2026
Builds on10
- High-Dimensional Sparse Linear BanditsBotao Hao, Tor Lattimore, Mengdi WangNeurIPS 2020 · 77 citations
- Partial Counterfactual Identification from Observational and Experimental DataJunzhe Zhang, Jin Tian, Elias BareinboimICML 2022 · 77 citations
- Causal Bandits with Unknown Graph StructureYangyi Lu, Amirhossein Meisami, Ambuj TewariNeurIPS 2021 · 53 citations
- General Transportability of Soft Interventions: Completeness ResultsJuan D. Correa, Elias BareinboimNeurIPS 2020 · 40 citations
- No Regrets for Learning the Prior in BanditsSoumya Basu, Branislav Kveton, Manzil Zaheer, Csaba SzepesváriNeurIPS 2021 · 39 citations
Related papers
- Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal ApproachWen Huang, Xintao WuAAAI 2024 · 2 citations
- Structural Causal Bandits under Markov EquivalenceMin Woo Park, Andy Arditi, Elias Bareinboim, Sanghack LeeNeurIPS 2025 · 3 citations
- Partial Structure Discovery is Sufficient for No-regret Learning in Causal BanditsMuhammad Qasim Elahi, Mahsa Ghasemi, Murat KocaogluNeurIPS 2024 · 11 citations
- Linear Causal Bandits: Unknown Graph and Soft InterventionsZirui Yan, Ali TajerNeurIPS 2024 · 11 citations
- Approximate Allocation Matching for Structural Causal Bandits with Unobserved ConfoundersLai Wei, Muhammad Qasim Elahi, Mahsa Ghasemi, Murat KocaogluNeurIPS 2023 · 10 citations
