An Orthogonal Learner for Individualized Outcomes in Markov Decision Processes
Emil Javurek, Valentyn Melnychuk, Jonas Schweisthal, Konstantin Hess, Dennis Frauen, Stefan Feuerriegel
摘要
Predicting individualized potential outcomes in sequential decision-making is central for optimizing therapeutic decisions in personalized medicine (e.g., which dosing sequence to give to a cancer patient). However, predicting potential outcomes over long horizons is notoriously difficult. Existing methods that break the curse of the horizon typically lack strong theoretical guarantees such as orthogonality and quasi-oracle efficiency. In this paper, we revisit the problem of predicting individualized potential outcomes in sequential decision-making (i.e., estimating Q-functions in Markov decision processes with observational data) through a causal inference lens. In particular, we develop a comprehensive theoretical foundation for meta-learners in this setting with a focus on beneficial theoretical properties. As a result, we yield a novel meta-learner called DRQ-learner and establish that it is: (1) doubly robust (i.e., valid inference under the misspecification of one of the models), (2) Neyman-orthogonal (i.e., insensitive to first-order estimation errors in the nuisance functions), and (3) achieves quasi-oracle efficiency (i.e., behaves asymptotically as if the ground-truth nuisance functions were known). Our DRQ-learner is applicable to settings with both discrete and continuous state spaces. Further, our DRQ-learner is flexible and can be used together with arbitrary machine learning models (e.g., neural networks). We validate our theoretical results through numerical experiments, thereby showing that our meta-learner outperforms state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- IGC-Net for conditional average potential outcome estimation over timeKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 8 次
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 5 次
它引用的顶会 Paper7
- Estimating counterfactual treatment outcomes over time through adversarially balanced representationsIoana Bica, Ahmed M. Alaa, James Jordon, Mihaela van der SchaarICLR 2020 · 被引用 224 次
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- Causal Transformer for Estimating Counterfactual OutcomesValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelICML 2022 · 被引用 146 次
- Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential EquationsNabeel Seedat, Fergus Imrie, Alexis Bellot, Zhaozhi Qian 等ICML 2022 · 被引用 68 次
- Deeply-Debiased Off-Policy Interval EstimationChengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui SongICML 2021 · 被引用 43 次
相关 Paper
- GDR-learners: Orthogonal Learning of Generative Models for Potential OutcomesValentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 1 次
- Conformal Meta-learners for Predictive Inference of Individual Treatment EffectsAhmed M. Alaa, Zaid Ahmad, Mark J. van der LaanNeurIPS 2023 · 被引用 32 次
- Reliable Off-Policy Learning for Dosage CombinationsJonas Schweisthal, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelNeurIPS 2023 · 被引用 22 次
- Rank-Learner: Orthogonal Ranking of Treatment EffectsHenri Arno, Dennis Frauen, Emil Javurek, Thomas Demeester 等ICML 2026
- Zero-shot causal learningHamed Nilforoshan, Michael Moor, Yusuf H. Roohani, Yining Chen 等NeurIPS 2023 · 被引用 25 次
