Does the Markov Decision Process Fit the Data: Testing for the Markov Property in Sequential Decision Making
Chengchun Shi, Runzhe Wan, Rui Song, Wenbin Lu, Ling Leng
2020年份
45被引次数
5顶会引用
摘要
The Markov assumption (MA) is fundamental to the empirical validity of reinforcement learning. In this paper, we propose a novel Forward-Backward Learning procedure to test MA in sequential decision making. The proposed test does not assume any parametric form on the joint distribution of the observed data and plays an important role for identifying the optimal policy in high-order Markov decision processes and partially observable MDPs. We apply our test to both synthetic datasets and a real data example from mobile health studies to illustrate its usefulness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Deeply-Debiased Off-Policy Interval EstimationChengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui SongICML 2021 · 被引用 43 次
- Data-Efficient Pipeline for Offline Reinforcement Learning with Limited DataAllen Nie, Yannis Flet-Berliac, Deon R. Jordan, William Steenbergen 等NeurIPS 2022 · 被引用 18 次
- Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement LearningShuguang Yu, Shuxing Fang, Ruixin Peng, Zhengling Qi 等NeurIPS 2024 · 被引用 9 次
- Simultaneous Statistical Inference for Off-Policy Evaluation in Reinforcement LearningTianpai Luo, Xinyuan Fan, Weichi WuNeurIPS 2025 · 被引用 1 次
- Off-Policy Evaluation under Nonignorable Missing DataHan Wang, Yang Xu, Wenbin Lu, Rui SongICML 2025
相关 Paper
- A Robust Test for the Stationarity Assumption in Sequential Decision MakingJitao Wang, Chengchun Shi, Zhenke WuICML 2023 · 被引用 9 次
- Reward Identification in Inverse Reinforcement LearningKuno Kim, Shivam Garg, Kirankumar Shiragur, Stefano ErmonICML 2021 · 被引用 43 次
- Asymptotically Optimal Sequential Testing with Markovian DataAlhad Sethi, SOFIA SAGAR KAVALI, Shubhada Agrawal, Debabrota Basu 等ICML 2026
- Robust Tests in Online Decision-MakingGi-Soo Kim, Jane P. Kim, Hyun-Joon YangAAAI 2022
- Finding Counterfactually Optimal Action Sequences in Continuous State SpacesStratis Tsirtsis, Manuel Gomez RodriguezNeurIPS 2023 · 被引用 18 次
