Preferential Temporal Difference Learning
Nishanth V. Anand, Doina Precup
Abstract
Temporal-Difference (TD) learning is a general and very useful tool for estimating the value function of a given policy, which in turn is required to find good policies. Generally speaking, TD learning updates states whenever they are visited. When the agent lands in a state, its value can be used to compute the TD-error, which is then propagated to other states. However, it may be interesting, when computing updates, to take into account other information than whether a state is visited or not. For example, some states might be more important than others (such as states which are frequently seen in a successful trajectory). Or, some states might have unreliable value estimates (for example, due to partial observability or lack of data), making their values less desirable as targets. We propose an approach to re-weighting states used in TD updates, both when they are the input and when they provide the target for the update. We prove that our approach converges with linear function approximation and illustrate its desirable empirical behaviour compared to other TD-style methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e5d6a06-e90b-453a-a8a9-7413bb8aed59Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Discerning Temporal Difference LearningJianfei MaAAAI 2024 · 2 citations
- Reducing Sampling Error in Batch Temporal Difference LearningBrahma S. Pavse, Ishan Durugkar, Josiah Hanna, Peter StoneICML 2020 · 14 citations
- Temporal Difference Learning as Gradient SplittingRui Liu, Alex OlshevskyICML 2021 · 18 citations
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman et al.NeurIPS 2023 · 17 citations
- Prediction and Control in Continual Reinforcement LearningNishanth Anand, Doina PrecupNeurIPS 2023 · 26 citations
