Lune

ICLR2026Top-tier venue

Q-learning with Posterior Sampling

Priyank Agrawal, Shipra Agrawal, Azmat Azati

2026Year
3Citations
1Top-tier citations

Abstract

Bayesian posterior sampling techniques have demonstrated superior empirical performance in many exploration-exploitation settings. However, their theoretical analysis remains a challenge, especially in complex settings like reinforcement learning. In this paper, we introduce Q-Learning with Posterior Sampling (PSQL), a simple Q-learning-based algorithm that uses Gaussian posteriors on Q-values for exploration, akin to the popular Thompson Sampling algorithm in the multi-armed bandit setting. We show that in the tabular episodic MDP setting, PSQL achieves a regret bound of O~(H2SAT)\tilde O(H^2\sqrt{SAT}), closely matching the known lower bound of Ω(HSAT)\Omega(H\sqrt{SAT}). Here, S, A denote the number of states and actions in the underlying Markov Decision Process (MDP), and T=KHT=KH with KK being the number of episodes and HH being the planning horizon. Our work provides several new technical insights into the core challenges in combining posterior sampling with dynamic programming and TD-learning-based RL algorithms, along with novel ideas for resolving those difficulties. We hope this will form a starting point for analyzing this efficient and important algorithmic technique in even more complex RL settings.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e562928c-e1d9-4f23-ba34-efc362f7958f

Cited by top-tier papers1

Ask how each one uses it

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines