An Orthogonal Learner for Individualized Outcomes in Markov Decision Processes
Emil Javurek, Valentyn Melnychuk, Jonas Schweisthal, Konstantin Hess, Dennis Frauen, Stefan Feuerriegel
Abstract
Predicting individualized potential outcomes in sequential decision-making is central for optimizing therapeutic decisions in personalized medicine (e.g., which dosing sequence to give to a cancer patient). However, predicting potential outcomes over long horizons is notoriously difficult. Existing methods that break the curse of the horizon typically lack strong theoretical guarantees such as orthogonality and quasi-oracle efficiency. In this paper, we revisit the problem of predicting individualized potential outcomes in sequential decision-making (i.e., estimating Q-functions in Markov decision processes with observational data) through a causal inference lens. In particular, we develop a comprehensive theoretical foundation for meta-learners in this setting with a focus on beneficial theoretical properties. As a result, we yield a novel meta-learner called DRQ-learner and establish that it is: (1) doubly robust (i.e., valid inference under the misspecification of one of the models), (2) Neyman-orthogonal (i.e., insensitive to first-order estimation errors in the nuisance functions), and (3) achieves quasi-oracle efficiency (i.e., behaves asymptotically as if the ground-truth nuisance functions were known). Our DRQ-learner is applicable to settings with both discrete and continuous state spaces. Further, our DRQ-learner is flexible and can be used together with arbitrary machine learning models (e.g., neural networks). We validate our theoretical results through numerical experiments, thereby showing that our meta-learner outperforms state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9eacfa0-30c9-4b6d-b116-7ad0e1f86bb1Cited by top-tier papers2
- IGC-Net for conditional average potential outcome estimation over timeKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 8 citations
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 5 citations
Builds on7
- Estimating counterfactual treatment outcomes over time through adversarially balanced representationsIoana Bica, Ahmed M. Alaa, James Jordon, Mihaela van der SchaarICLR 2020 · 224 citations
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- Causal Transformer for Estimating Counterfactual OutcomesValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelICML 2022 · 146 citations
- Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential EquationsNabeel Seedat, Fergus Imrie, Alexis Bellot, Zhaozhi Qian et al.ICML 2022 · 68 citations
- Deeply-Debiased Off-Policy Interval EstimationChengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui SongICML 2021 · 43 citations
Related papers
- GDR-learners: Orthogonal Learning of Generative Models for Potential OutcomesValentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 1 citation
- Conformal Meta-learners for Predictive Inference of Individual Treatment EffectsAhmed M. Alaa, Zaid Ahmad, Mark J. van der LaanNeurIPS 2023 · 32 citations
- Reliable Off-Policy Learning for Dosage CombinationsJonas Schweisthal, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelNeurIPS 2023 · 22 citations
- Rank-Learner: Orthogonal Ranking of Treatment EffectsHenri Arno, Dennis Frauen, Emil Javurek, Thomas Demeester et al.ICML 2026
- Zero-shot causal learningHamed Nilforoshan, Michael Moor, Yusuf H. Roohani, Yining Chen et al.NeurIPS 2023 · 25 citations
