Lune

NeurIPS2024Top-tier venue

Finding good policies in average-reward Markov Decision Processes without prior knowledge

Adrienne Tuynman, Rémy Degenne, Emilie Kaufmann

2024Year
14Citations
4Top-tier citations

Abstract

We revisit the identification of an ε\varepsilon-optimal policy in average-reward Markov Decision Processes (MDP). In such MDPs, two measures of complexity have appeared in the literature: the diameter, DD, and the optimal bias span, HH, which satisfy H≤DH\leq D. Prior work have studied the complexity of ε\varepsilon-optimal policy identification only when a generative model is available. In this case, it is known that there exists an MDP with D≃HD \simeq H for which the sample complexity to output an ε\varepsilon-optimal policy is Ω(SAD/ε2)\Omega(SAD/\varepsilon^2) where SS and AA are the sizes of the state and action spaces. Recently, an algorithm with a sample complexity of order SAH/ε2SAH/\varepsilon^2 has been proposed, but it requires the knowledge of HH. We first show that the sample complexity required to estimate HH is not bounded by any function of S,AS,A and HH, ruling out the possibility to easily make the previous algorithm agnostic to HH. By relying instead on a diameter estimation procedure, we propose the first algorithm for (ε,δ)(\varepsilon,\delta)-PAC policy identification that does not need any form of prior knowledge on the MDP. Its sample complexity scales in SAD/ε2SAD/\varepsilon^2 in the regime of small ε\varepsilon, which is near-optimal. In the online setting, our first contribution is a lower bound which implies that a sample complexity polynomial in HH cannot be achieved in this setting. Then, we propose an online algorithm with a sample complexity in SAD2/ε2SAD^2/\varepsilon^2, as well as a novel approach based on a data-dependent stopping rule that we believe is promising to further reduce this bound.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 8534d9df-3d58-4a1b-b43e-6ffe5b53eaf9

Cited by top-tier papers4

Ask how each one uses it

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines