A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and Preconditioning
Zechen Wu, Amy Greenwald, Ronald E. Parr
Abstract
In off-policy policy evaluation (OPE) tasks within reinforcement learning, Temporal Difference Learning(TD) and Fitted Q-Iteration (FQI) have traditionally been viewed as differing in the number of updates toward the target value function: TD makes one update, FQI makes an infinite number, and Partial Fitted Q-Iteration (PFQI) performs a finite number. We show that this view is not accurate, and provide a new mathematical perspective under linear value function approximation that unifies these methods as a single iterative method solving the same linear system, but using different matrix splitting schemes and preconditioners. We show that increasing the number of updates under the same target value function, i.e., the target network technique, is a transition from using a constant preconditioner to using a data-feature adaptive preconditioner. This elucidates, for the first time, why TD convergence does not necessarily imply FQI convergence, and establishes tight convergence connections among TD, PFQI, and FQI. Our framework enables sharper theoretical results than previous work and characterization of the convergence conditions for each algorithm, without relying on assumptions about the features (e.g., linear independence). We also provide an encoder-decoder perspective to better understand the convergence conditions of TD, and prove, for the first time, that when a large learning rate doesn't work, trying a smaller one may help. Our framework also leads to the discovery of new crucial conditions on features for convergence, and shows how common assumptions about features influence convergence, e.g., the assumption of linearly independent features can be dropped without compromising the convergence guarantees of stochastic TD in the on-policy setting. This paper is also the first to introduce matrix splitting into the convergence analysis of these algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cab10f38-4b7a-4d48-8536-222ed3add6d7Builds on4
- Representations for Stable Off-Policy Reinforcement LearningDibya Ghosh, Marc G. BellemareICML 2020 · 46 citations
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman et al.NeurIPS 2023 · 17 citations
- Why Target Networks Stabilise Temporal Difference MethodsMattie Fellows, Matthew J. A. Smith, Shimon WhitesonICML 2023 · 10 citations
- Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function ApproximationFengdi Che, Chenjun Xiao, Jincheng Mei, Bo Dai et al.ICML 2024 · 7 citations
Related papers
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
- Understanding and Leveraging Overparameterization in Recursive Value EstimationChenjun Xiao, Bo Dai, Jincheng Mei, Oscar A. Ramirez et al.ICLR 2022 · 17 citations
- Temporal Difference Learning as Gradient SplittingRui Liu, Alex OlshevskyICML 2021 · 18 citations
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
