Parameterized Indexed Value Function for Efficient Exploration in Reinforcement Learning
Tian Tan, Zhihan Xiong, Vikranth R. Dwaracherla
Abstract
It is well known that quantifying uncertainty in the action-value estimates is crucial for efficient exploration in reinforcement learning. Ensemble sampling offers a relatively computationally tractable way of doing this using randomized value functions. However, it still requires a huge amount of computational resources for complex problems. In this paper, we present an alternative, computationally efficient way to induce exploration using index sampling. We use an indexed value function to represent uncertainty in our action-value estimates. We first present an algorithm to learn parameterized indexed value function through a distributional version of temporal difference in a tabular setting and prove its regret bound. Then, in a computational point of view, we propose a dual-network architecture, Parameterized Indexed Networks (PINs), comprising one mean network and one uncertainty network to learn the indexed value function. Finally, we show the efficacy of PINs through computational experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a54172e2-6b55-486e-9d2e-1ea9590cc8feRelated papers
- Universal Value-Function UncertaintiesMoritz Akiya Zanger, Max Weltevrede, Yaniv Oren, Pascal R. van der Vaart et al.ICLR 2026 · 1 citation
- Hypermodels for ExplorationVikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Ian Osband et al.ICLR 2020 · 49 citations
- Exploration via Epistemic Value EstimationSimon Schmitt, John Shawe-Taylor, Hado van HasseltAAAI 2023 · 4 citations
- Model-Free Active Exploration in Reinforcement LearningAlessio Russo, Alexandre ProutièreNeurIPS 2023 · 7 citations
- Deep Bandits Show-Off: Simple and Efficient Exploration with Deep NetworksRong Zhu, Mattia RigottiNeurIPS 2021 · 10 citations
