Lune

ICLR2026顶会

Instance-Dependent Fixed-Budget Pure Exploration in Reinforcement Learning

Yeongjong Kim, Yeoneung Kim, Kwang-Sung Jun

出版方
2026年份

摘要

We study the problem of fixed budget pure exploration in reinforcement learning.The goal is to identify a near-optimal policy, given a fixed budget on the number of interactions with the environment. Unlike the standard PAC setting, we do not require the target error level ϵ\epsilon and failure rate δ\delta as input. We propose novel algorithms and provide, to the best of our knowledge, the first instance-dependent ϵ\epsilon-uniform guarantee, meaning that the probability that ϵ\epsilon-correctness is ensured can be obtained simultaneously for all ϵ\epsilon above a budget-dependent threshold. It characterizes the budget requirements in terms of the problem-specific hardness of exploration. As a core component of our analysis, we derive a ϵ\epsilon-uniform guarantee for the multiple bandit problem—solving multiple multi-armed bandit instances simultaneously—which may be of independent interest. To enable our analysis, we also develop tools for reward-free exploration under the fixed-budget setting, which we believe will be useful for future work.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖