Learning Decision Policies with Instrumental Variables through Double Machine Learning
Daqian Shao, Ashkan Soleymani, Francesco Quinzan, Marta Kwiatkowska
Abstract
A common issue in learning decision-making policies in data-rich settings is spurious correlations in the offline dataset, which can be caused by hidden confounders. Instrumental variable (IV) regression, which utilises a key unconfounded variable known as the instrument, is a standard technique for learning causal relationships between confounded action, outcome, and context variables. Most recent IV regression algorithms use a two-stage approach, where a deep neural network (DNN) estimator learnt in the first stage is directly plugged into the second stage, in which another DNN is used to estimate the causal effect. Naively plugging the estimator can cause heavy bias in the second stage, especially when regularisation bias is present in the first stage estimator. We propose DML-IV, a non-linear IV regression method that reduces the bias in two-stage IV regressions and effectively learns high-performing policies. We derive a novel learning objective to reduce bias and design the DML-IV algorithm following the double/debiased machine learning (DML) framework. The learnt DML-IV estimator has strong convergence rate and suboptimality guarantees that match those when the dataset is unconfounded. DML-IV outperforms state-of-the-art IV regression methods on IV regression benchmarks and learns high-performing policies in the presence of instruments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 142ed56f-570c-4391-b56d-e372d5c0973dCited by top-tier papers2
- Demystifying Spectral Feature Learning for Instrumental Variable RegressionDimitri Meunier, Antoine Moulin, Jakub Wornbard, Vladimir Kostic et al.NeurIPS 2025 · 5 citations
- Causal Imitation Learning under Expert-Observable and Expert-Unobservable ConfoundingDaqian Shao, Thomas Kleine Buening, Marta KwiatkowskaICLR 2026 · 1 citation
Builds on14
- Learning Counterfactual Representations for Estimating Individual Dose-Response CurvesPatrick Schwab, Lorenz Linhardt, Stefan Bauer, Joachim M. Buhmann et al.AAAI 2020 · 159 citations
- Estimating the Effects of Continuous-valued Interventions using Generative Adversarial NetworksIoana Bica, James Jordon, Mihaela van der SchaarNeurIPS 2020 · 137 citations
- Dual Instrumental Variable RegressionKrikamol Muandet, Arash Mehrjou, Si Kai Lee, Anant RajNeurIPS 2020 · 87 citations
- Learning Deep Features in Instrumental Variable RegressionLiyuan Xu, Yutian Chen, Siddarth Srinivasan, Nando de Freitas et al.ICLR 2021 · 85 citations
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 81 citations
Related papers
- Conditional Instrumental Variable Regression with Representation Learning for Causal InferenceDebo Cheng, Ziqi Xu, Jiuyong Li, Lin Liu et al.ICLR 2024 · 14 citations
- Optimality and Adaptivity of Deep Neural Features for Instrumental Variable RegressionJuno Kim, Dimitri Meunier, Arthur Gretton, Taiji Suzuki et al.ICLR 2025
- Nonparametric Instrumental Variable Regression through Stochastic Approximate GradientsYuri R. Fonseca, Caio Peixoto, Yuri F. SaporitoNeurIPS 2024 · 8 citations
- Estimating individual treatment effects under unobserved confounding using binary instrumentsDennis Frauen, Stefan FeuerriegelICLR 2023 · 3 citations
- Causal Inference with Conditional Instruments Using Deep Generative ModelsDebo Cheng, Ziqi Xu, Jiuyong Li, Lin Liu et al.AAAI 2023 · 24 citations
