Lune

ICLR2026顶会

Learning in Prophet Inequalities with Noisy Observations

Jung-hun Kim, Vianney Perchet

2026年份

摘要

We study the prophet inequality, a fundamental problem in online decision-making and optimal stopping, in a practical setting where rewards are observed only through noisy realizations and reward distributions are unknown. At each stage, the decision-maker receives a noisy reward whose true value follows a linear model with an unknown latent parameter, and observes a feature vector drawn from a distribution. To address this challenge, we propose algorithms that integrate learning and decision-making via lower-confidence-bound (LCB) thresholding. In the i.i.d. setting, we establish that both an Explore-then-Decide strategy and an ε\varepsilon-Greedy variant achieve the sharp competitive ratio of 1−1/e1 - 1/e, under a mild condition on the optimal value. For non-identical distributions, we show that a competitive ratio of 1/21/2 can be guaranteed against a relaxed benchmark. Moreover, with limited window access to past rewards, the tight ratio of 1/21/2 against the optimal benchmark is achieved.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 9a18fa8d-b0f2-41a4-a59f-e80872d8731e

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖