Off-Policy Evaluation with Policy-Dependent Optimization Response
Wenshuo Guo, Michael I. Jordan, Angela Zhou
Abstract
The intersection of causal inference and machine learning for decision-making is rapidly expanding, but the default decision criterion remains an average of individual causal outcomes across a population. In practice, various operational restrictions ensure that a decision-maker's utility is not realized as an average but rather as an output of a downstream decision-making problem (such as matching, assignment, network flow, minimizing predictive risk). In this work, we develop a new framework for off-policy evaluation with policy-dependent linear optimization responses: causal outcomes introduce stochasticity in objective function coefficients. Under this framework, a decision-maker's utility depends on the policy-dependent optimization, which introduces a fundamental challenge of optimization bias even for the case of policy evaluation. We construct unbiased estimators for the policy-dependent estimand by a perturbation method, and discuss asymptotic variance properties for a set of adjusted plug-in estimators. Lastly, attaining unbiased policy evaluation allows for policy optimization: we provide a general algorithm for optimizing causal interventions. We corroborate our theoretical results with numerical simulations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Decision-Focused Learning with Directional GradientsMichael Huang, Vishal GuptaNeurIPS 2024 · 25 citations
- Empirical Gateaux Derivatives for Causal InferenceMichael I. Jordan, Yixin Wang, Angela ZhouNeurIPS 2022 · 13 citations
Builds on6
- Performative PredictionJuan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz HardtICML 2020 · 422 citations
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller et al.NeurIPS 2021 · 64 citations
- RieszNet and ForestRiesz: Automatic Debiased Machine Learning with Neural Nets and Random ForestsVictor Chernozhukov, Whitney Newey, Victor Quintas-Martinez, Vasilis SyrgkanisICML 2022 · 61 citations
- Interference, Bias, and Variance in Two-Sided Marketplace Experimentation: Guidance for PlatformsHannah Li, Geng Zhao, Ramesh Johari, Gabriel Y. WeintraubWWW 2022 · 46 citations
- Cost-Effective Incentive Allocation via Structured Counterfactual InferenceRomain Lopez, Chenchen Li, Xiang Yan, Junwu Xiong et al.AAAI 2020 · 21 citations
Related papers
- Causal Strategic Linear RegressionYonadav Shavit, Benjamin L. Edelman, Brian AxelrodICML 2020 · 91 citations
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 5 citations
- Strategic Instrumental Variable Regression: Recovering Causal Relationships From Strategic ResponsesKeegan Harris, Dung Daniel T. Ngo, Logan Stapleton, Hoda Heidari et al.ICML 2022 · 37 citations
- Causal Modeling for Fairness In Dynamical SystemsElliot Creager, David Madras, Toniann Pitassi, Richard S. ZemelICML 2020 · 72 citations
- Estimation of Bounds on Potential Outcomes For Decision MakingMaggie Makar, Fredrik D. Johansson, John V. Guttag, David A. SontagICML 2020 · 11 citations
