Deciding What to Learn: A Rate-Distortion Approach
Dilip Arumugam, Benjamin Van Roy
Abstract
Agents that learn to select optimal actions represent a prominent focus of the sequential decision-making literature. In the face of a complex environment or constraints on time and resources, however, aiming to synthesize such an optimal policy can become infeasible. These scenarios give rise to an important trade-off between the information an agent must acquire to learn and the sub-optimality of the resulting policy. While an agent designer has a preference for how this trade-off is resolved, existing approaches further require that the designer translate these preferences into a fixed learning target for the agent. In this work, leveraging rate-distortion theory, we automate this process such that the designer need only express their preferences via a single hyperparameter and the agent is endowed with the ability to compute its own learning targets that best achieve the desired trade-off. We establish a general bound on expected discounted regret for an agent that decides what to learn in this manner along with computational experiments that illustrate the expressiveness of designer preferences and even show improvements over Thompson sampling in identifying an optimal policy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6be8f2bd-4f21-4290-bce2-7d349eb6735eCited by top-tier papers6
- Efficient Exploration for LLMsVikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, Benjamin Van RoyICML 2024 · 45 citations
- Deciding What to Model: Value-Equivalent Sampling for Reinforcement LearningDilip Arumugam, Benjamin Van RoyNeurIPS 2022 · 25 citations
- The Value of Information When Deciding What to LearnDilip Arumugam, Benjamin Van RoyNeurIPS 2021 · 19 citations
- Domain Generalization without Excess Empirical RiskOzan Sener, Vladlen KoltunNeurIPS 2022 · 10 citations
- SIA: Symbolic Interpretability for Anticipatory Deep Reinforcement Learning in Network ControlMohammadErfan Jabbari, Abhishek Duttagupta, Claudio Fiandrino, Leonardo Bonati et al.INFOCOM 2026 · 1 citation
Builds on2
Related papers
- On Bits and Bandits: Quantifying the Regret-Information Trade-offItai Shufaro, Nadav Merlis, Nir Weinberger, Shie MannorICLR 2025
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths et al.NeurIPS 2022 · 30 citations
- Sequential Transfer in Reinforcement Learning with a Generative ModelAndrea Tirinzoni, Riccardo Poiani, Marcello RestelliICML 2020 · 26 citations
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 2 citations
- Making RL with Preference-based Feedback Efficient via RandomizationRunzhe Wu, Wen SunICLR 2024 · 44 citations
