Efficient Policy Learning from Surrogate-Loss Classification Reductions
Andrew Bennett, Nathan Kallus
Abstract
Recent work on policy learning from observational data has highlighted the importance of efficient policy evaluation and has proposed reductions to weighted (cost-sensitive) classification. But, efficient policy evaluation need not yield efficient estimation of policy parameters. We consider the estimation problem given by a weighted surrogate-loss classification reduction of policy learning with any score function, either direct, inverse-propensity weighted, or doubly robust. We show that, under a correct specification assumption, the weighted classification formulation need not be efficient for policy parameters. We draw a contrast to actual (possibly weighted) binary classification, where correct specification implies a parametric model, while for policy learning it only implies a semiparametric model. In light of this, we instead propose an estimation approach based on generalized method of moments, which is efficient for the policy parameters. We propose a particular method based on recent developments on solving moment problems using neural networks and demonstrate the efficiency and regret benefits of this method empirically.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0b714e3-9d3f-4d6c-95f6-5f14b9b13dc8Cited by top-tier papers8
- Reliable Off-Policy Learning for Dosage CombinationsJonas Schweisthal, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelNeurIPS 2023 · 22 citations
- Fair Off-Policy Learning from Observational DataDennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICML 2024 · 11 citations
- Functional Generalized Empirical Likelihood Estimation for Conditional Moment RestrictionsHeiner Kremer, Jia-Jie Zhu, Krikamol Muandet, Bernhard SchölkopfICML 2022 · 9 citations
- Treatment Effect Estimation for Optimal Decision-MakingDennis Frauen, Valentyn Melnychuk, Jonas Schweisthal, Mihaela van der Schaar et al.NeurIPS 2025 · 8 citations
- Estimation Beyond Data Reweighting: Kernel Method of MomentsHeiner Kremer, Yassine Nemmour, Bernhard Schölkopf, Jia-Jie ZhuICML 2023 · 7 citations
Related papers
- Off-policy estimation with adaptively collected data: the power of online learningJeonghwan Lee, Cong MaNeurIPS 2024 · 4 citations
- Learning from Biased Data: A Semi-Parametric ApproachPatrice Bertail, Stéphan Clémençon, Yannick Guyonvarch, Nathan NoiryICML 2021 · 6 citations
- Doubly Robust Distributionally Robust Offline Contextual PricingMin Xu, Xinyi Yin, Yunfan Zhang, Yuxuan Han et al.ICML 2026 · 15 citations
- Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningNathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan ZhouICML 2022 · 39 citations
- A Variational Framework for Estimating Continuous Treatment Effects with Measurement ErrorErdun Gao, Howard D. Bondell, Wei Huang, Mingming GongICLR 2024 · 7 citations
