Optimizing Sequential Experimental Design with Deep Reinforcement Learning
Tom Blau, Edwin V. Bonilla, Iadine Chades, Amir Dezfouli
Abstract
Bayesian approaches developed to solve the optimal design of sequential experiments are mathematically elegant but computationally challenging. Recently, techniques using amortization have been proposed to make these Bayesian approaches practical, by training a parameterized policy that proposes designs efficiently at deployment time. However, these methods may not sufficiently explore the design space, require access to a differentiable probabilistic model and can only optimize over continuous design spaces. Here, we address these limitations by showing that the problem of optimizing policies can be reduced to solving a Markov decision process (MDP). We solve the equivalent MDP with modern deep reinforcement learning techniques. Our experiments show that our approach is also computationally efficient at deployment time and exhibits state-of-the-art performance on both continuous and discrete design spaces, even when the probabilistic model is a black box. 1 CSIRO's Data61, Australia 2 CSIRO's Land and Water, Australia.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1625bd3-8e06-4a1e-9ea9-2a435f532434Cited by top-tier papers24
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental DesignDeepro Choudhury, Sinead Williamson, Adam Golinski, Ning Miao et al.ICLR 2026 · 24 citations
- Amortized Bayesian Experimental Design for Decision-MakingDaolang Huang, Yujia Guo, Luigi Acerbi, Samuel KaskiNeurIPS 2024 · 24 citations
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Differentiable Multi-Target Causal Bayesian Experimental DesignPanagiotis Tigas, Yashas Annadani, Desi R. Ivanova, Andrew Jesson et al.ICML 2023 · 15 citations
- Amortized Active Causal Induction with Deep Reinforcement LearningYashas Annadani, Panagiotis Tigas, Stefan Bauer, Adam FosterNeurIPS 2024 · 13 citations
Builds on7
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Deep Adaptive Design: Amortizing Sequential Bayesian Experimental DesignAdam Foster, Desi R. Ivanova, Ilyas Malik, Tom RainforthICML 2021 · 119 citations
- Bayesian Experimental Design for Implicit Models by Mutual Information Neural EstimationSteven Kleinegesse, Michael U. GutmannICML 2020 · 84 citations
- Implicit Deep Adaptive Design: Policy-Based Experimental Design without LikelihoodsDesi R. Ivanova, Adam Foster, Steven Kleinegesse, Michael U. Gutmann et al.NeurIPS 2021 · 81 citations
- Uncertainty-aware Active Learning for Optimal Bayesian ClassifierGuang Zhao, Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander et al.ICLR 2021 · 43 citations
Related papers
- Nesting Particle Filters for Experimental Design in Dynamical SystemsSahel Iqbal, Adrien Corenflos, Simo Särkkä, Hany AbdulsamadICML 2024 · 10 citations
- Constrained Bayesian Experimental Design via Online PlanningYujia Guo, Daolang Huang, Xinyu Zhang, Sammie Katt et al.ICML 2026 · 1 citation
- Step-DAD: Semi-Amortized Policy-Based Bayesian Experimental DesignMarcel Hedman, Desi R. Ivanova, Cong Guan, Tom RainforthICML 2025
- Bayesian Optimization over Discrete and Mixed Spaces via Probabilistic ReparameterizationSamuel Daulton, Xingchen Wan, David Eriksson, Maximilian Balandat et al.NeurIPS 2022 · 71 citations
- PABBO: Preferential Amortized Black-Box OptimizationXinyu Zhang, Daolang Huang, Samuel Kaski, Julien MartinelliICLR 2025
