A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and Preconditioning
Zechen Wu, Amy Greenwald, Ronald E. Parr
摘要
In off-policy policy evaluation (OPE) tasks within reinforcement learning, Temporal Difference Learning(TD) and Fitted Q-Iteration (FQI) have traditionally been viewed as differing in the number of updates toward the target value function: TD makes one update, FQI makes an infinite number, and Partial Fitted Q-Iteration (PFQI) performs a finite number. We show that this view is not accurate, and provide a new mathematical perspective under linear value function approximation that unifies these methods as a single iterative method solving the same linear system, but using different matrix splitting schemes and preconditioners. We show that increasing the number of updates under the same target value function, i.e., the target network technique, is a transition from using a constant preconditioner to using a data-feature adaptive preconditioner. This elucidates, for the first time, why TD convergence does not necessarily imply FQI convergence, and establishes tight convergence connections among TD, PFQI, and FQI. Our framework enables sharper theoretical results than previous work and characterization of the convergence conditions for each algorithm, without relying on assumptions about the features (e.g., linear independence). We also provide an encoder-decoder perspective to better understand the convergence conditions of TD, and prove, for the first time, that when a large learning rate doesn't work, trying a smaller one may help. Our framework also leads to the discovery of new crucial conditions on features for convergence, and shows how common assumptions about features influence convergence, e.g., the assumption of linearly independent features can be dropped without compromising the convergence guarantees of stochastic TD in the on-policy setting. This paper is also the first to introduce matrix splitting into the convergence analysis of these algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Representations for Stable Off-Policy Reinforcement LearningDibya Ghosh, Marc G. BellemareICML 2020 · 被引用 46 次
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman 等NeurIPS 2023 · 被引用 17 次
- Why Target Networks Stabilise Temporal Difference MethodsMattie Fellows, Matthew J. A. Smith, Shimon WhitesonICML 2023 · 被引用 10 次
- Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function ApproximationFengdi Che, Chenjun Xiao, Jincheng Mei, Bo Dai 等ICML 2024 · 被引用 7 次
相关 Paper
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 被引用 61 次
- Understanding and Leveraging Overparameterization in Recursive Value EstimationChenjun Xiao, Bo Dai, Jincheng Mei, Oscar A. Ramirez 等ICLR 2022 · 被引用 17 次
- Temporal Difference Learning as Gradient SplittingRui Liu, Alex OlshevskyICML 2021 · 被引用 18 次
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
