Lune

ICLR2026顶会

A Unifying View of Coverage in Linear Off-policy Evaluation

Philip Amortila, Audrey Huang, Akshay Krishnamurthy, Nan Jiang

2026年份
2被引次数

摘要

Off-policy evaluation (OPE) is a fundamental task in reinforcement learning (RL). In the classic setting of linear OPE, finite-sample guarantees often take the form

Prediction error≤poly(Cπ,d,1/n,log(1/δ)),\textrm{Prediction error} \le \textrm{poly}(C^\pi, d, 1/n, log(1/\delta)),

where dd is the dimension of the features, and CπC^\pi is a feature coverage parameter that characterizes the degree to which the visited features lie in the span of the data distribution. While such guarantees are well-understood for several popular algorithms under stronger assumptions (e.g. Bellman completeness), the understanding is lacking and fragmented in the minimal setting where the target value function is linearly realizable in the features. Despite recent interest in tight characterizations of the statistical rate in this setting, the right notion of coverage remains unclear, and candidate definitions from prior analyses have undesirable properties and are starkly disconnected from more standard definitions in the literature.

We provide a novel finite-sample analysis of a canonical algorithm for this setting, LSTDQ. Inspired by an instrumental-variable view, we develop error bounds that depend on a novel coverage parameter, the feature-dynamics coverage, which can be interpreted as linear coverage in an induced dynamical system for feature evolution. With further assumptions---such as Bellman-completeness---our definition successfully recovers the coverage parameters specialized to those settings, finally yielding a unified understanding for coverage in linear OPE.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext c81f0eba-e519-4d5e-a4c3-7cb365fa774a

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖