Deciding What to Learn: A Rate-Distortion Approach
Dilip Arumugam, Benjamin Van Roy
摘要
Agents that learn to select optimal actions represent a prominent focus of the sequential decision-making literature. In the face of a complex environment or constraints on time and resources, however, aiming to synthesize such an optimal policy can become infeasible. These scenarios give rise to an important trade-off between the information an agent must acquire to learn and the sub-optimality of the resulting policy. While an agent designer has a preference for how this trade-off is resolved, existing approaches further require that the designer translate these preferences into a fixed learning target for the agent. In this work, leveraging rate-distortion theory, we automate this process such that the designer need only express their preferences via a single hyperparameter and the agent is endowed with the ability to compute its own learning targets that best achieve the desired trade-off. We establish a general bound on expected discounted regret for an agent that decides what to learn in this manner along with computational experiments that illustrate the expressiveness of designer preferences and even show improvements over Thompson sampling in identifying an optimal policy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Efficient Exploration for LLMsVikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, Benjamin Van RoyICML 2024 · 被引用 45 次
- Deciding What to Model: Value-Equivalent Sampling for Reinforcement LearningDilip Arumugam, Benjamin Van RoyNeurIPS 2022 · 被引用 25 次
- The Value of Information When Deciding What to LearnDilip Arumugam, Benjamin Van RoyNeurIPS 2021 · 被引用 19 次
- Domain Generalization without Excess Empirical RiskOzan Sener, Vladlen KoltunNeurIPS 2022 · 被引用 10 次
- SIA: Symbolic Interpretability for Anticipatory Deep Reinforcement Learning in Network ControlMohammadErfan Jabbari, Abhishek Duttagupta, Claudio Fiandrino, Leonardo Bonati 等INFOCOM 2026 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- On Bits and Bandits: Quantifying the Regret-Information Trade-offItai Shufaro, Nadav Merlis, Nir Weinberger, Shie MannorICLR 2025
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths 等NeurIPS 2022 · 被引用 30 次
- Sequential Transfer in Reinforcement Learning with a Generative ModelAndrea Tirinzoni, Riccardo Poiani, Marcello RestelliICML 2020 · 被引用 26 次
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 被引用 2 次
- Making RL with Preference-based Feedback Efficient via RandomizationRunzhe Wu, Wen SunICLR 2024 · 被引用 44 次
